Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 526–600

1321 questions total · 18pages · All types, answers revealed

Page 7

Page 8 of 18

Page 9
526
MCQeasy

A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is in CSV format and is updated daily. The engineer wants to use a fully managed service that can handle the load without managing infrastructure. Which AWS service should the engineer use?

A.AWS Glue
B.AWS Database Migration Service (AWS DMS)
C.Amazon Redshift Spectrum
D.Amazon Kinesis Data Firehose
AnswerA

AWS Glue is a fully managed extract, transform, and load (ETL) service that can read from S3 and write to Redshift. It handles provisioning and scaling of the underlying infrastructure. The engineer can create a Glue job to load the CSV data into Redshift on a schedule. This meets the requirement for a managed service without infrastructure management.

Why this answer

AWS Glue is a fully managed ETL service that can read CSV files from Amazon S3, apply transformations if needed, and load the data into Amazon Redshift. It handles the underlying infrastructure, scales automatically, and can be scheduled to run daily. This makes it the ideal choice for the scenario.

Exam trap

The trap here is confusing Redshift Spectrum with an ingestion service; Spectrum queries data in S3 but does not load it into Redshift, so it does not meet the ingestion requirement.

527
MCQeasy

A company is using Amazon S3 to store log files. The security team requires that all data be encrypted in transit. Which of the following ensures encryption in transit for S3?

A.Use HTTPS (SSL/TLS) when accessing S3 endpoints.
B.Use Amazon S3 Transfer Acceleration.
C.Enable client-side encryption before uploading to S3.
D.Use server-side encryption with S3 managed keys (SSE-S3).
AnswerA

HTTPS (SSL/TLS) encrypts data between the client and S3 endpoints, directly satisfying the in-transit encryption requirement. Server-side encryption options such as SSE-S3 or SSE-KMS protect data at rest only, leaving network traffic exposed. Enforcing TLS via bucket policies that deny non-secure transport further guarantees this.

Why this answer

Encryption in transit for Amazon S3 is ensured by using HTTPS (SSL/TLS) when accessing S3 endpoints. HTTPS encrypts data as it travels between the client and S3, protecting it from eavesdropping. This is the standard method to enforce encryption in transit for S3, and it can be enforced via bucket policies that deny requests not using HTTPS.

Exam trap

DEA-C01 often tests the confusion between encryption at rest and encryption in transit, leading candidates to select server-side encryption options that only protect data at rest.

How to eliminate wrong answers

Option B is wrong because Amazon S3 Transfer Acceleration speeds up uploads by using AWS edge locations, but it does not provide encryption in transit; it still relies on HTTPS for encryption. Option C is wrong because client-side encryption encrypts data at rest before uploading, but it does not encrypt data in transit; the data is already encrypted when it leaves the client, but the transmission itself may not be secure unless HTTPS is used. Option D is wrong because server-side encryption with S3 managed keys (SSE-S3) encrypts data at rest, not in transit; it does not protect data during transmission.

528
MCQmedium

A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains inconsistent date formats in a column named 'transaction_date'. The engineer needs to standardize all dates to ISO 8601 format (YYYY-MM-DD) and then write the cleaned data to a new S3 location. Which transformation should the engineer apply?

A.Use the 'Change data type' transformation to convert the column to a date type.
B.Use the 'Replace value' transformation to replace all date separators with hyphens.
C.Use the 'Format date' transformation with a custom date format pattern.
D.Use the 'Split column' transformation to separate day, month, and year into individual columns.
AnswerC

The 'Format date' transformation in AWS Glue DataBrew allows specifying input and output date formats. By setting the output format to 'yyyy-MM-dd', the engineer can standardize all dates to ISO 8601. This transformation handles multiple input formats if configured correctly, making it the appropriate choice for this scenario.

Why this answer

The 'Format date' transformation in AWS Glue DataBrew is designed to parse and reformat date values. By specifying the input format(s) and setting the output format to 'yyyy-MM-dd', the engineer can standardize inconsistent date strings into ISO 8601. This ensures consistency for downstream processing and analytics, and it handles multiple input patterns without manual string manipulation.

Exam trap

The trap here is confusing data type conversion with date formatting; changing the type does not reformat the string, and simple string replacements cannot handle varying date orders.

529
MCQeasy

A company stores its application logs in Amazon S3. The logs are generated daily and need to be retained for 3 years for compliance. The logs are accessed frequently for the first 30 days, occasionally for the next 6 months, and rarely after that. The data engineering team wants to minimize storage costs while ensuring that logs are available for retrieval within 12 hours for the first 6 months and within 48 hours after that. The team also wants to automatically delete logs after 3 years. Which lifecycle policy should the team implement?

A.Transition to S3 One Zone-IA after 30 days, delete after 6 months.
B.Transition to S3 Standard-IA after 30 days, delete after 3 years.
C.Transition to S3 Standard-IA after 30 days, to S3 Glacier after 6 months, delete after 3 years.
D.Transition to S3 Standard-IA after 30 days, to S3 Glacier Deep Archive after 6 months, delete after 3 years.
AnswerD

Meets cost and retrieval time requirements.

Why this answer

It aligns with the access patterns and retrieval requirements: S3 Standard-IA after 30 days for occasional access with immediate retrieval, then S3 Glacier Deep Archive after 6 months for rare access with a 12-hour retrieval time (via expedited or standard retrieval), and deletion after 3 years. This minimizes storage costs while meeting the 48-hour retrieval window for older logs.

Exam trap

The trap here is that candidates may choose Option C (S3 Glacier) thinking it is the cheapest cold storage, but S3 Glacier Deep Archive is actually the lowest-cost option for data that is rarely accessed and can tolerate a 12-hour retrieval time, which still satisfies the 48-hour requirement.

How to eliminate wrong answers

Option A is wrong because it transitions to S3 One Zone-IA after 30 days, which does not provide the durability or availability needed for compliance logs, and it deletes after 6 months instead of 3 years. Option B is wrong because it keeps logs in S3 Standard-IA for the entire 3 years, which is more expensive than transitioning to a colder storage class after 6 months, and it does not meet the cost-minimization goal. Option C is wrong because it transitions to S3 Glacier after 6 months, which has a retrieval time of 1-5 minutes for expedited or 3-5 hours for standard, exceeding the 48-hour requirement but not being the most cost-effective option; S3 Glacier Deep Archive is cheaper and still meets the 48-hour retrieval window.

530
MCQeasy

A data engineer is using AWS Glue to run a job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job must be scheduled to run every day at 2:00 AM UTC. Which AWS Glue feature should the engineer use to set up this schedule?

A.AWS Glue Triggers with a schedule type of 'Scheduled' and a cron expression.
B.AWS Glue job bookmarks with a time-based configuration.
C.AWS Glue DataBrew schedules with a cron expression.
D.AWS Glue Workflows with an on-demand trigger started by an Amazon EventBridge rule.
AnswerA

Glue Triggers support scheduled triggers defined with cron expressions in UTC. The engineer can create a trigger with a cron expression like 'cron(0 2 * * ? *)' to run the job daily at 2:00 AM UTC. This is the native Glue scheduling mechanism and integrates directly with the job, requiring no external orchestration.

Why this answer

AWS Glue Triggers of type 'Scheduled' accept cron expressions and run jobs on a defined schedule in UTC. Creating a scheduled trigger with the appropriate cron expression is the direct, native way to run a Glue job daily at 2:00 AM UTC without external orchestration.

Exam trap

The trap here is confusing Glue job bookmarks, which track processed data, with Glue triggers, which actually schedule job runs.

531
MCQhard

A data engineer is using Amazon Redshift and needs to improve query performance for a workload that involves frequent joins between a large fact table and a small dimension table. The dimension table is updated daily with new records. The engineer wants to minimize data movement during joins. Which Redshift distribution style should the engineer use for the dimension table?

A.AUTO
B.EVEN
C.KEY
D.ALL
AnswerD

ALL distribution replicates the entire dimension table to every compute node. This ensures that joins with the fact table can be performed locally on each node without data movement, significantly improving join performance. It is ideal for small, slowly changing dimension tables.

Why this answer

ALL distribution replicates the entire dimension table to every node, eliminating the need to redistribute data during joins with the fact table. This minimizes network traffic and speeds up join operations. It is best suited for small dimension tables that are frequently joined with larger tables.

Exam trap

The trap here is assuming that KEY distribution on the join column always eliminates data movement, but it does not when the fact table is distributed differently.

532
MCQhard

A data engineer is designing a data lake on Amazon S3. The compliance team requires that objects be automatically deleted after 7 years. Additionally, objects must be transitioned to Amazon S3 Glacier Instant Retrieval after 30 days to reduce costs. Which S3 lifecycle policy configuration meets these requirements?

A.Transition to Glacier Instant Retrieval after 30 days, then expire after 90 days.
B.Transition to Glacier Instant Retrieval after 30 days, then expire after 2555 days.
C.Transition to Glacier Deep Archive after 30 days, then expire after 7 years.
D.Transition to S3 Standard-IA after 30 days, then expire after 7 years.
AnswerB

The lifecycle rule transitions objects to Glacier Instant Retrieval at day 30, meeting the cost requirement, and expires them at day 2555, which equals seven years. Both the transition and retention constraints in the stem are satisfied by this single rule.

Why this answer

The requirement is to transition objects to Glacier Instant Retrieval after 30 days and delete them after 7 years. 7 years equals 2555 days (7 × 365). Option B correctly specifies transition after 30 days and expiration after 2555 days. This meets both the cost optimization and compliance retention requirements.

Exam trap

The trap is miscalculating 7 years in days (2555) or confusing Glacier Instant Retrieval with Glacier Deep Archive, leading to wrong storage class or expiration values.

How to eliminate wrong answers

Option A is wrong because expiring after 90 days does not meet the 7-year retention requirement. Option C is wrong because it transitions to Glacier Deep Archive instead of Glacier Instant Retrieval, and Deep Archive has retrieval times of hours, not instant. Option D is wrong because it transitions to S3 Standard-IA instead of Glacier Instant Retrieval, which does not meet the specified storage class requirement.

533
MCQeasy

A company wants to import data from an external FTP server into Amazon S3 on a daily basis. The data volumes are moderate. Which AWS service is MOST suitable for this task?

A.Amazon S3 Transfer Acceleration
B.AWS Transfer Family
C.AWS DataSync
D.AWS Glue with a JDBC connection
AnswerB

AWS Transfer Family provides a managed FTP endpoint that ingests files directly into Amazon S3, satisfying the daily external FTP import requirement without custom scripts or servers. It supports scheduled or continuous transfers and handles the moderate volume comfortably.

Why this answer

AWS Transfer Family is the most suitable service because it provides fully managed support for SFTP, FTPS, and FTP protocols, enabling direct, secure file transfers from an external FTP server to Amazon S3 without needing to manage any infrastructure. It integrates natively with S3, so files are automatically stored in the specified bucket upon transfer completion, making it ideal for daily imports of moderate data volumes.

Exam trap

The trap here is that candidates often confuse AWS DataSync with a general-purpose file transfer tool, but DataSync requires an agent on the source and does not natively support FTP protocols, whereas AWS Transfer Family is purpose-built for FTP-based transfers to S3.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Transfer Acceleration is a feature that speeds up uploads to S3 over the internet by using AWS edge locations, but it does not support FTP protocols or act as a server-side endpoint to receive files from an external FTP server. Option C is wrong because AWS DataSync is designed for moving large volumes of data between on-premises storage and AWS services, but it requires installing an agent on the source environment and does not natively support FTP as a source protocol. Option D is wrong because AWS Glue with a JDBC connection is intended for extracting data from databases using JDBC drivers, not for handling file transfers over FTP/SFTP/FTPS protocols.

534
MCQmedium

A data engineer is building a real-time data pipeline to ingest sensor data from IoT devices. The data is sent to AWS IoT Core, which publishes messages to a Kinesis Data Stream. Each message is about 1 KB in size. The data must be transformed (add a device location field) and then stored in Amazon S3 for long-term analytics. The engineer has set up a Lambda function to transform the records and write to S3. However, the engineer notices that the Lambda function is invoked thousands of times per second, causing high costs and occasional throttling. The Lambda function processes only one record at a time. The engineer wants to reduce the number of Lambda invocations and improve throughput. What should the engineer do?

A.Reduce the number of shards in the Kinesis stream to limit concurrency.
B.Increase the Lambda function's memory allocation to improve performance.
C.Replace the Lambda function with Amazon Kinesis Data Firehose and use its built-in transformation.
D.Configure the event source mapping to use a larger batch size and set a batch window.
AnswerD

Increasing the batch size lets each Lambda invocation process many records instead of one, and the batch window accumulates records before invoking. Both settings cut invocation counts and throttling while raising throughput, satisfying the cost and performance goals.

Why this answer

Configuring the event source mapping with a larger batch size and a batch window allows Lambda to process multiple records per invocation, reducing the number of invocations and costs. This improves throughput and reduces throttling. Option A is incorrect because reducing shards reduces the stream capacity, causing backpressure and potential data loss.

Option B is incorrect because increasing memory does not reduce the number of invocations; it only speeds up processing per invocation, but still processes one record at a time. Option C is incorrect because Kinesis Data Firehose can batch records, but it still uses per-record Lambda transformation if you use a Lambda function, or it can use built-in transformations but not the flexible logic described. The most direct solution is to batch records in the existing Lambda function via event source mapping parameters.

Exam trap

A candidate might think that reducing the number of shards will reduce invocations, but that actually reduces the stream's ability to handle the data volume and can cause throttling or data loss.

535
Multi-Selectmedium

Which TWO actions should a data engineer take to optimize Amazon S3 query performance for Amazon Athena when dealing with large Parquet files? (Choose 2.)

Select 2 answers
A.Store data in a single large file without partitioning
B.Use GZIP compression on the Parquet files
C.Split large files into many small files
D.Optimize file sizes to be around 64 MB to 256 MB
E.Partition the data by frequently filtered columns
AnswersD, E

Athena splits Parquet scans across files; many small files add per-file overhead, while very large files limit parallelism. Targeting 64–256 MB balances scan parallelism against request overhead, directly improving query performance on large Parquet datasets in S3.

Why this answer

Option D is correct because Athena's performance degrades with very large Parquet files since a single file can only be read by one split; targeting roughly 64 MB to 256 MB per file allows Athena's split-based parallel reads to distribute work across more nodes. Option E is correct because partitioning data by frequently filtered columns (e.g., date, region) enables partition pruning, so Athena scans only the relevant S3 prefixes instead of the entire dataset, drastically reducing bytes scanned and query time. Option A is wrong because a single large unpartitioned file forces full-table scans and prevents parallelism.

Option B is wrong because Parquet already uses columnar compression internally, and GZIP is not splittable, which can actually hurt parallelism for large files. Option C is wrong because splitting into many tiny files creates excessive per-file overhead (metadata, open/close costs) and degrades Athena performance rather than improving it.

536
MCQhard

A data engineer needs to share a dataset stored in an S3 bucket with a partner AWS account. The partner should be able to read the data without needing to authenticate with the engineer's account. The engineer must not share any secret keys. Which approach should be used?

A.Write a bucket policy that grants access to the partner account's IAM role.
B.Generate presigned URLs and share them with the partner.
C.Make the bucket publicly readable.
D.Create an IAM user with access keys and share them with the partner.
AnswerA

A bucket policy granting the partner account's IAM role read access satisfies the cross-account, credential-free requirement: the partner's principals assume their own roles, and AWS evaluates the resource-based policy, so no secret keys or authentication in the engineer's account are needed.

Why this answer

A bucket policy granting read access to the partner account's IAM role is the correct approach because it enables cross-account access without sharing credentials. The partner's IAM role assumes its own identity, and the bucket policy explicitly authorizes that principal, so no secret keys are exchanged. This is the standard AWS cross-account access pattern for S3 data sharing.

Exam trap

DEA-C01 often tests whether candidates default to presigned URLs or public buckets for cross-account sharing, when the correct pattern is a bucket policy granting the partner principal access without credential sharing.

How to eliminate wrong answers

Option B is wrong because presigned URLs are time-limited and tied to the signer's credentials; they are unsuitable for ongoing dataset access and expire, requiring regeneration. Option C is wrong because making the bucket publicly readable exposes data to the entire internet, violating least privilege and likely compliance requirements. Option D is wrong because sharing IAM user access keys violates AWS best practices, cannot be audited per-user, and creates long-lived credentials that are hard to rotate.

537
MCQmedium

A data engineer needs to store semi-structured JSON data from IoT devices. The data is written once, read rarely, but must be queryable using SQL. The storage cost must be minimized. Which storage solution should the engineer choose?

A.Store JSON in Amazon Redshift as SUPER data type
B.Store JSON in an Amazon RDS for MySQL table
C.Store JSON documents in Amazon DynamoDB and use PartiQL for queries
D.Store JSON files in Amazon S3 and use Amazon Athena for queries
AnswerD

S3 provides the lowest-cost durable storage for write-once, rarely read JSON, and Athena queries that data in place using standard SQL over the AWS Glue Data Catalog, avoiding the cost of a continuously running database or cluster.

Why this answer

Amazon S3 provides the lowest-cost storage for data that is written once and rarely read, while Amazon Athena enables serverless SQL querying directly on JSON files stored in S3. This combination minimizes storage costs because S3 charges only for the data stored and retrieval, with no minimum fees or provisioning required, and Athena charges only for the data scanned per query. The workload's write-once, read-rarely pattern aligns perfectly with S3's durability and lifecycle policies, making it the most cost-effective choice.

Exam trap

The trap here is that candidates often choose DynamoDB (Option C) because it supports JSON natively and PartiQL provides SQL-like queries, but they overlook that DynamoDB's provisioned throughput and storage costs are significantly higher than S3 for write-once, read-rarely workloads, and that Athena on S3 is the serverless, cost-optimized solution for ad-hoc SQL queries on infrequently accessed data.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for high-performance analytics on structured and semi-structured data, but it incurs significant costs for provisioned clusters even when data is rarely queried, making it unsuitable for minimizing storage costs. Option B is wrong because Amazon RDS for MySQL is a relational database that requires provisioning and paying for a database instance 24/7, and storing JSON in a MySQL table incurs overhead for indexing and transactions that are unnecessary for write-once, read-rarely data. Option C is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for low-latency, high-throughput workloads, but its storage costs are higher than S3 for rarely accessed data, and PartiQL queries on DynamoDB still consume read capacity units, leading to ongoing costs that exceed S3+Athena for infrequent queries.

538
MCQhard

Refer to the exhibit. A data engineer is reviewing the configuration of an Amazon Redshift cluster. The engineer wants to ensure that the cluster can be restored to a point in time up to 35 days in the past. Based on the exhibit, what change is needed?

A.Increase the automated snapshot retention period to 35 days.
B.Change the cluster subnet group to a custom one.
C.Enable encryption on the cluster.
D.Increase the number of nodes to 6.
AnswerA

Automated snapshots are the mechanism that supports point-in-time restoration in Amazon Redshift. The retention period defaults to one day and can be extended to a maximum of 35 days, so raising it to 35 days directly satisfies the stated requirement to restore to any point within that window.

Why this answer

Amazon Redshift automated snapshots have a default retention period of 1 day, but can be configured up to 35 days. To restore to a point in time up to 35 days, the retention period must be increased to 35 days.

Exam trap

DEA-C01 often tests the default and maximum snapshot retention periods, leading candidates to assume the default is already 35 days or that other cluster settings affect retention.

How to eliminate wrong answers

Option B is wrong because the cluster subnet group affects network placement, not snapshot retention. Option C is wrong because encryption is for data security, not backup retention. Option D is wrong because the number of nodes affects performance and storage, not snapshot retention.

539
Multi-Selectmedium

A data engineer is configuring an Amazon Redshift cluster and needs to optimize query performance for complex analytical queries that involve large joins. The engineer wants to reduce the amount of data movement during query execution. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the number of nodes in the cluster to add more compute resources.
B.Define sort keys on columns frequently used in join and filter conditions.
C.Enable concurrency scaling to handle concurrent queries.
D.Choose a distribution style that colocates joined tables on the same node slices.
E.Use columnar storage for all tables.
AnswersB, D

Sort keys determine the order in which data is stored on disk. When sort keys are defined on columns used in joins and filters, Redshift can use zone maps to skip irrelevant blocks and perform efficient merge joins. This reduces I/O and data movement, enhancing query performance for complex analytical workloads.

Why this answer

To reduce data movement during joins in Amazon Redshift, the engineer should choose a distribution style that colocates joined tables on the same node slices and define sort keys on columns used in joins and filters. These actions minimize network traffic and enable efficient merge joins. Increasing nodes, enabling concurrency scaling, or using columnar storage do not directly address data movement within a query.

Exam trap

The trap here is assuming that adding more nodes or enabling concurrency scaling will reduce data movement during joins, but those actions address capacity and concurrency, not join efficiency.

540
Multi-Selecthard

Which THREE are valid considerations when troubleshooting data loss in an AWS Glue ETL job? (Choose three.)

Select 3 answers
A.Job bookmarks may be skipping new data if not configured properly.
B.Server-side encryption is disabled on the S3 bucket.
C.The job timeout is set too low.
D.Dynamic frame transformations may drop rows with errors.
E.The mapping of source columns to target columns may be incorrect.
AnswersA, D, E

Job bookmarks track previously processed data using a state store. If misconfigured — for example, wrong transformation context or a changed source path — the bookmark may mark new files or rows as already processed, silently skipping them and producing apparent data loss in the Glue ETL output.

Why this answer

Option A is correct because AWS Glue job bookmarks track previously processed data; if a bookmark is misconfigured, stale, or reset, the job can skip new or changed records, causing apparent data loss. Option D is correct because Glue DynamicFrame transformations and write operations can drop or quarantine rows that fail schema or type validation, so records with errors may silently disappear unless error handling is configured. Option E is correct because an incorrect ApplyMapping or column mapping can write source columns into the wrong target fields or omit them, making data appear lost in the target.

Option B is not a valid consideration because S3 server-side encryption protects data at rest and does not cause records to be dropped during ETL processing. Option C is not a valid consideration because a low job timeout causes the job to fail or stop, which is a job failure rather than silent data loss.

Exam trap

DEA-C01 often tests whether candidates can distinguish data-loss causes from job-failure or security issues — options like low timeout or disabled encryption sound plausible but cause failures or posture problems, not silent data loss.

541
Multi-Selecteasy

A data engineer is setting up a data pipeline using Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using an AWS Lambda function before delivery. Which THREE steps are required to configure this?

Select 3 answers
A.Create a Lambda@Edge function in the same Region.
B.Create an AWS Lambda function that transforms the data.
C.Attach an IAM role to the Firehose delivery stream that grants permission to invoke the Lambda function.
D.Configure an S3 event notification to trigger the Lambda function when new data arrives.
E.Configure the Kinesis Data Firehose delivery stream to use the Lambda function as a data transformation source.
AnswersB, C, E

Firehose requires a Lambda function to perform record transformation; the function must exist and be selected in the delivery stream's processing configuration. Creating it is therefore a mandatory step before Firehose can invoke it on incoming records.

Why this answer

Option B is correct because Firehose data transformation requires an AWS Lambda function (a standard regional Lambda function, not Lambda@Edge) that receives batches of records and returns transformed records in the expected Firehose format. Option C is correct because the Firehose delivery stream assumes an IAM role, and that role must include lambda:InvokeFunction permission on the transformation function so Firehose can call it. Option E is correct because you must enable and select the Lambda function in the Firehose delivery stream's processing configuration (the Lambda transformation processor) so records are transformed before being written to Amazon S3.

Option A is wrong because Lambda@Edge runs at CloudFront edge locations for viewer/origin request and response events, not for Firehose record transformation. Option D is wrong because S3 event notifications trigger actions after objects land in S3; Firehose invokes the transformation Lambda itself, so no S3 event notification is needed.

Exam trap

DEA-C01 often tests the misconception that S3 event notifications are needed to trigger Lambda for Firehose transformation, but Firehose directly invokes Lambda; candidates must remember the direct integration.

542
Multi-Selecthard

A data engineering team is building a data lake on Amazon S3. They need to ingest data from multiple sources: (1) streaming IoT data, (2) daily CSV exports from an on-premises system via SFTP, and (3) change data capture (CDC) from an Amazon Aurora database. Which THREE services should the team use to ingest these data sources?

Select 3 answers
A.Amazon Kinesis Data Streams for IoT data ingestion.
B.AWS Database Migration Service (DMS) for CDC from Aurora.
C.AWS Transfer Family for SFTP-based file ingestion.
D.AWS Glue ETL for CDC from Aurora.
E.Amazon EMR for daily CSV ingestion.
AnswersA, B, C

Kinesis Data Streams ingests continuous, high-volume IoT telemetry with durable ordering and replay, matching the streaming source. It feeds downstream consumers that land the data in Amazon S3, satisfying the real-time ingestion requirement of the data lake.

Why this answer

Amazon Kinesis Data Streams (A) is the right choice for streaming IoT data because it is a fully managed, scalable service designed for real-time ingestion of high-volume streaming records into S3 via Kinesis Data Firehose or consumers. AWS Database Migration Service (B) supports ongoing change data capture (CDC) from Amazon Aurora (and other relational databases), replicating inserts, updates, and deletes continuously, which matches the CDC requirement. AWS Transfer Family (C) provides a fully managed SFTP endpoint backed by Amazon S3, so the on-premises system can push daily CSV exports over SFTP directly into the data lake.

AWS Glue ETL (D) is a batch/transform service, not a CDC ingestion mechanism, so it does not satisfy the Aurora CDC requirement. Amazon EMR (E) is a big data processing cluster, not a managed ingestion service for daily SFTP CSV files, so it is not the appropriate ingestion tool here.

Exam trap

The trap here is confusing AWS Glue ETL (a batch ETL tool) with AWS DMS (a database migration and CDC service), and assuming Amazon EMR is an ingestion service rather than a processing framework for large-scale data transformations.

543
MCQmedium

Refer to the exhibit. A data engineer has attached this IAM policy to an AWS Glue job role. The Glue job fails when trying to write transformed data to an S3 bucket located in a different AWS account. What is the most likely reason?

A.The policy does not allow lambda:InvokeAsync
B.The Glue job role does not have permissions to write to S3
C.The policy does not grant s3:ListBucket, and the bucket policy may not allow cross-account access
D.The policy does not include kinesis:DescribeStream
AnswerC

Cross-account S3 access requires both bucket policy and IAM permissions, including ListBucket.

Why this answer

The IAM policy shown does not include the s3:ListBucket permission, which is required for the Glue job to list objects in the S3 bucket before writing. Additionally, cross-account access requires both the source account's IAM policy (this one) to grant write permissions and the target account's S3 bucket policy to explicitly allow the source account's role, which may not be configured. Without s3:ListBucket, the Glue job cannot verify the bucket's existence or structure, causing the write operation to fail.

Exam trap

The trap here is that candidates assume s3:PutObject alone is sufficient for writing to S3, but AWS requires s3:ListBucket for bucket-level operations like listing and validation, especially in cross-account scenarios where the bucket's existence must be confirmed.

How to eliminate wrong answers

Option A is wrong because lambda:InvokeAsync is a permission for invoking AWS Lambda functions asynchronously, which is irrelevant to writing data to S3 from a Glue job. Option B is wrong because the policy does include s3:PutObject and s3:PutObjectAcl, which are write permissions; the failure is due to missing s3:ListBucket, not a lack of write permissions entirely. Option D is wrong because kinesis:DescribeStream is a permission for Amazon Kinesis streams, which is unrelated to S3 write operations in this cross-account scenario.

544
MCQmedium

A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Aurora MySQL. After the migration, the data in Aurora is inconsistent with the source. The engineer needs to ensure ongoing replication with minimal downtime. Which solution should the engineer implement?

A.Use AWS Schema Conversion Tool (SCT) to convert the schema
B.Export the data from Oracle and import into Aurora using mysqldump
C.Configure a DMS task with change data capture (CDC)
D.Perform a full load migration again
AnswerC

Change data capture reads the source Oracle redo logs continuously, replicating only committed inserts, updates and deletes to Aurora MySQL after the initial full load. This satisfies the ongoing replication requirement with minimal downtime, since no bulk re-copy or application outage is needed to keep the target consistent.

Why this answer

Configuring a DMS task with Change Data Capture (CDC) enables ongoing replication of changes from the source Oracle database to the target Aurora MySQL database with minimal downtime, ensuring consistency. Option A is incorrect because AWS Schema Conversion Tool (SCT) only converts schema and does not handle data replication. Option B is incorrect because mysqldump provides a one-time export/import, not ongoing replication.

Option D is incorrect because performing a full load migration again would disrupt operations and would not capture ongoing changes.

545
MCQeasy

A data engineer needs to run a PySpark transformation on a 2 TB dataset stored in Amazon S3 and write the output back to S3 in Parquet. The team wants to use AWS Glue but does not want to manage clusters or tune Spark configuration manually. They also want to pay only for the time the job runs. Which AWS Glue component should the engineer use?

A.AWS Glue DataBrew recipe job
B.Amazon EMR on EC2 cluster with a bootstrap action
C.AWS Glue Python shell job
D.AWS Glue for Apache Spark job
AnswerD

AWS Glue for Apache Spark is a serverless Spark environment. You supply a PySpark script, and Glue provisions and manages the Spark cluster, applies tuning defaults, and bills per DPU-hour for the job's runtime. This matches the requirement for custom PySpark on a 2 TB S3 dataset without cluster management and with pay-per-run pricing.

Why this answer

AWS Glue for Apache Spark provides a serverless, distributed Spark runtime. The engineer submits a PySpark script, Glue handles cluster provisioning and Spark tuning, and billing is based on DPU-hours consumed during the run. This fits the need for custom PySpark on a 2 TB S3 dataset with no cluster management and pay-per-run cost, unlike single-node shell jobs or self-managed EMR clusters.

Exam trap

The trap here is equating 'AWS Glue' generically with any Glue job type, when Python shell jobs cannot run distributed PySpark and DataBrew targets visual preparation rather than custom code.

546
MCQeasy

A company is using an Amazon RDS for PostgreSQL database to store application data. The data engineering team needs to run complex analytical queries that join multiple large tables. These queries are causing performance degradation on the production database. The team wants to offload the analytical workload to a separate system that can handle large-scale data processing. Which AWS service should the team use?

A.Amazon Redshift
B.Amazon ElastiCache for Redis
C.Amazon RDS for MySQL
D.Amazon DynamoDB
AnswerA

Amazon Redshift is a fully managed, petabyte-scale data warehouse designed for analytical queries. It uses columnar storage and massively parallel processing to handle complex joins and aggregations efficiently. Offloading analytical workloads to Redshift frees up the RDS production database and provides better performance for large-scale data processing.

Why this answer

Amazon Redshift is purpose-built for analytical workloads, offering columnar storage, massively parallel processing, and advanced query optimization. By moving the analytical queries to Redshift, the team can achieve faster performance without impacting the production RDS database. This separation of transactional and analytical workloads is a common best practice.

Exam trap

The trap here is assuming that any database can handle analytical queries, but transactional databases like RDS are not optimized for large-scale joins and aggregations.

547
MCQmedium

A company is using Amazon DynamoDB to store session data for a web application. The data engineer needs to ensure that the data is encrypted at rest. Which action should the data engineer take?

A.Enable encryption at rest on the DynamoDB Accelerator (DAX) cluster.
B.Use client-side encryption before writing to DynamoDB.
C.Ensure encryption at rest is enabled on the DynamoDB table (default).
D.Enable DynamoDB Time to Live (TTL) to encrypt data.
AnswerC

DynamoDB encrypts all data at rest by default using AWS-owned keys, so no action is needed to enable it. Verifying that default encryption is active on the table satisfies the requirement without configuring client-side or KMS key options.

Why this answer

DynamoDB tables are encrypted at rest by default using AWS Key Management Service (KMS) with an AWS owned key. The data engineer does not need to take any additional action to enable encryption at rest, as it is automatically enabled for all new DynamoDB tables. This ensures that all data stored on disk, including the table's primary key, local secondary indexes, and global secondary indexes, is encrypted before being written to SSDs in AWS data centers.

Exam trap

The trap here is that candidates may assume encryption at rest must be explicitly enabled or configured, when in fact DynamoDB enables it by default for all tables, leading them to incorrectly select client-side encryption or other unrelated options.

How to eliminate wrong answers

Option A is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that does not store data at rest; it only caches data in memory, and encryption at rest is not a configurable feature for DAX clusters. Option B is wrong because client-side encryption is an additional security measure that encrypts data before sending it to DynamoDB, but it is not required to meet the requirement of encryption at rest, which is already provided by DynamoDB's default server-side encryption. Option D is wrong because DynamoDB Time to Live (TTL) is a feature that automatically deletes expired items from a table to manage storage costs; it does not provide any encryption functionality.

548
MCQhard

A data engineer is designing a data lake on Amazon S3 and needs to ensure that data is encrypted at rest. The company requires that encryption keys be managed by AWS and automatically rotated annually. The engineer also wants to audit key usage. Which S3 encryption option should the engineer choose?

A.Server-side encryption with Amazon S3 managed keys (SSE-S3).
B.Server-side encryption with customer-provided keys (SSE-C).
C.Client-side encryption with a custom key management solution.
D.Server-side encryption with AWS KMS managed keys (SSE-KMS).
AnswerD

SSE-KMS uses AWS KMS to manage encryption keys, which can be automatically rotated annually. AWS KMS integrates with AWS CloudTrail to log key usage, providing the required auditability. This option meets both the encryption and auditing requirements.

Why this answer

SSE-KMS uses AWS KMS to manage encryption keys, supports automatic annual key rotation, and logs key usage in AWS CloudTrail. This provides both the required encryption and auditability. Other options either do not support AWS-managed rotation or lack auditing capabilities.

Exam trap

The trap here is assuming that SSE-S3 provides auditing because it is AWS-managed, but only KMS-backed keys provide CloudTrail logs of key usage.

549
MCQeasy

A data engineer needs to audit all AWS KMS key usage events for the past 90 days to verify compliance. Which AWS service should be used?

A.VPC Flow Logs
B.AWS CloudTrail
C.AWS Config
D.Amazon Inspector
AnswerB

AWS CloudTrail records KMS API calls such as Encrypt, Decrypt and GenerateDataKey, capturing the identity, timestamp and key involved. This gives the 90-day audit trail of key usage events needed to verify compliance, which KMS alone does not provide.

Why this answer

(AWS CloudTrail). AWS CloudTrail logs all API calls made to AWS KMS, including key usage events, and retains event history for the past 90 days by default, making it suitable for auditing. Option A (VPC Flow Logs) captures network traffic, not API calls.

Option C (AWS Config) tracks resource configuration changes, not API calls. Option D (Amazon Inspector) performs vulnerability assessments, not API logging.

550
Multi-Selectmedium

A data engineer is designing a data ingestion pipeline using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format and must be converted to Apache Parquet before storage. The engineer wants to minimize costs and operational effort. Which two actions should the engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable compression on the Firehose delivery stream to reduce storage costs.
B.Enable record format conversion in the Firehose delivery stream and select Apache Parquet as the output format.
C.Use AWS Glue to create a table in the AWS Glue Data Catalog that defines the schema of the JSON data for the Firehose stream to reference during conversion.
D.Configure the Firehose delivery stream to invoke an AWS Lambda function to convert each record from JSON to Parquet.
E.Set the Firehose buffer size to the maximum value to reduce the number of S3 PUT operations.
AnswersB, C

Firehose supports record format conversion from JSON to Parquet using an AWS Glue table that defines the schema. This built-in feature eliminates the need for custom processing, reducing operational effort and cost. It is the recommended approach for format conversion in Firehose.

Why this answer

To convert JSON to Parquet in Firehose with minimal effort, enable record format conversion and create an AWS Glue table defining the schema. These two actions allow Firehose to perform the conversion automatically. Other options either add overhead, do not address the conversion, or are supplementary but not required.

Exam trap

The trap here is thinking that a Lambda function is needed for format conversion, but Firehose has native support using AWS Glue Data Catalog.

551
MCQmedium

A data engineer is setting up cross-account access to an encrypted S3 bucket. The bucket uses a customer-managed KMS key. The engineer has configured the bucket policy and the IAM role in the source account. The target account still gets access denied errors when trying to read objects. What is the most likely cause?

A.The KMS key policy does not grant the target account's IAM role the kms:Decrypt permission.
B.The S3 bucket has Object Ownership set to BucketOwnerPreferred.
C.The bucket policy does not allow the target account's root user.
D.The VPC Endpoint policy blocks access from the target account.
AnswerA

S3 access to KMS-encrypted objects requires permissions on both the bucket policy and the KMS key policy. Because the key is customer-managed, its key policy must explicitly grant the target account's IAM role kms:Decrypt; otherwise KMS denies the request despite correct IAM and bucket configuration.

Why this answer

For cross-account access to an S3 bucket encrypted with a customer-managed KMS key, the KMS key policy must explicitly grant the target account's IAM role the kms:Decrypt permission. Without this, the target account will get access denied errors even if the bucket policy and IAM role are correctly configured. Option B is incorrect because Object Ownership does not affect cross-account read access.

Option C is incorrect because the bucket policy only needs to allow the target account's IAM role, not the root user. Option D is incorrect because VPC Endpoint policies are not relevant to this cross-account access issue.

552
MCQeasy

A company stores application logs in Amazon S3 and needs to query them using standard SQL. The logs are in JSON format and are updated daily. The data engineering team wants a serverless solution that requires minimal management and can automatically discover the schema. Which AWS service should they use?

A.Amazon EMR with Apache Hive.
B.Amazon Athena with an AWS Glue Data Catalog table.
C.Amazon RDS for PostgreSQL with the aws_s3 extension.
D.Amazon Redshift Spectrum.
AnswerB

Amazon Athena is a serverless query service that allows SQL queries directly on S3 data. By using an AWS Glue Data Catalog table, the schema can be automatically discovered and updated. This combination requires no infrastructure management and is ideal for querying JSON logs stored in S3 with minimal overhead.

Why this answer

Amazon Athena is a serverless query service that integrates with the AWS Glue Data Catalog to automatically discover schemas and query data directly in S3. It requires no infrastructure management and supports standard SQL, making it ideal for querying JSON logs with minimal operational overhead.

Exam trap

The trap here is confusing serverless querying with managed cluster services like Redshift Spectrum or EMR, which require more management and are not fully serverless.

553
MCQhard

A data engineer is monitoring an AWS Glue job that reads from an Amazon S3 bucket and writes to Amazon Redshift. The job has been running for 2 hours, which is longer than usual. The engineer checks the Glue job's metrics and sees that the number of active executors is high, but the job is not making progress. The engineer suspects a data skew issue. Which action should the engineer take to diagnose and mitigate the skew?

A.Use the Spark UI to inspect the stage details and identify tasks that are taking significantly longer than others.
B.Increase the number of DPUs for the Glue job to add more executors and reduce the impact of skew.
C.Enable AWS Glue job metrics and examine the 'glue.driver.aggregate.bytesRead' metric to identify skewed partitions.
D.Modify the Glue job to use the 'groupFiles' option to combine small files and improve parallelism.
AnswerA

The Spark UI provides detailed information about stages and tasks, including task duration and input size. By examining the stage details, you can identify tasks that are processing much more data than others, indicating data skew. This is the most direct way to diagnose skew in a Glue job.

Why this answer

Data skew in Spark jobs manifests as a few tasks taking much longer than others. The Spark UI is the primary tool to diagnose skew by showing task-level metrics. Once identified, mitigation strategies include salting keys, using broadcast joins, or repartitioning.

The correct action here is to use the Spark UI to inspect stage details.

Exam trap

The trap here is assuming that adding more resources (DPUs) will solve skew, when in fact skew requires redistributing data, not just adding compute.

554
Multi-Selectmedium

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to ensure that the job handles data quality issues such as duplicate records and missing values before loading. The job must also minimize the amount of data shuffled across the network. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Use the FillMissingValues transform to replace missing values with a default.
B.Repartition the DynamicFrame to a single partition before writing to Redshift.
C.Use the DropDuplicates transform to remove duplicate records.
D.Use the ResolveChoice transform to handle data type conflicts.
E.Enable job bookmarks to track processed data and avoid reprocessing.
AnswersA, C

FillMissingValues is an AWS Glue transform that fills null or missing values in specified columns with a default value. This handles missing values as required. It is a built-in transform that can be easily added to the job, ensuring data quality before loading to Redshift, and it does not require custom code.

Why this answer

DropDuplicates removes duplicate records, and FillMissingValues replaces missing values with defaults, directly addressing the data quality requirements. Both are built-in AWS Glue transforms that can be applied without custom code. The other options do not handle duplicates or missing values, and repartitioning to a single partition would increase network shuffle, contrary to the requirement.

Exam trap

The trap here is confusing data quality transforms with performance tuning; repartitioning and job bookmarks are often mistakenly associated with data cleaning but do not resolve duplicates or missing values.

555
MCQhard

A data engineer stores raw customer records in an Amazon S3 bucket and runs an AWS Glue job that writes curated Parquet files to a second bucket. The governance team requires that the curated data carry a verifiable record of which job run produced it and that any modification to a curated file be detectable. The engineer must also prove that the curated dataset has not been altered since a nightly baseline. Which combination of AWS features should the engineer use?

A.Enable AWS CloudTrail data events on the curated bucket and use AWS Glue job bookmarks to record the last processed run in the Data Catalog.
B.Compute an SHA-256 checksum per curated file, store the checksums and Glue job run IDs in a manifest, and use that manifest to verify the files against the nightly baseline.
C.Enable S3 Versioning on the curated bucket and configure an S3 event notification to write each version ID to an Amazon DynamoDB table keyed by the Glue job run ID.
D.Enable S3 Object Lock in governance mode on the curated bucket and attach a retention period equal to the nightly baseline interval.
AnswerB

A per-file SHA-256 checksum recorded alongside the Glue job run ID gives both provenance and tamper evidence. Recomputing checksums and comparing them with the stored manifest detects any byte-level change to a curated file, and the run ID ties each file to the job execution that produced it, which is exactly what the governance team requires.

Why this answer

Provenance and tamper evidence require linking each curated file to the job run that created it and being able to prove the bytes have not changed. A manifest that pairs a Glue job run ID with a per-file SHA-256 checksum delivers both. Comparing the nightly recomputed checksums against the baseline manifest surfaces any alteration, while versioning, CloudTrail, and Object Lock address adjacent concerns but not content integrity.

Exam trap

The trap here is treating S3 Versioning or CloudTrail logging as proof of integrity, when only a content digest compared against a trusted baseline can reveal that file bytes were altered.

556
MCQeasy

A data engineer needs to store transaction data that requires strong consistency, ACID transactions, and complex join queries. Which AWS service is most appropriate?

A.Amazon DynamoDB
B.Amazon RDS for PostgreSQL
C.Amazon S3
D.Amazon Redshift
AnswerB

Amazon RDS for PostgreSQL provides ACID transactions, strong consistency and full SQL joins, matching all three stated requirements. Its relational engine supports complex multi-table joins natively, unlike purpose-built NoSQL stores that trade joins and cross-item transactions for scale.

Why this answer

Amazon RDS for PostgreSQL is the most appropriate choice because it provides full ACID transaction support, strong consistency, and the ability to perform complex join queries using standard SQL. Unlike NoSQL or data warehouse solutions, PostgreSQL is a relational database that excels at enforcing referential integrity and supporting multi-table joins with advanced indexing.

Exam trap

The trap here is that candidates often confuse DynamoDB's 'eventually consistent reads' with strong consistency, or assume its limited transaction API can replace full ACID relational databases, but the question explicitly requires complex joins and ACID transactions, which only a relational database like PostgreSQL can provide.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database that does not support complex join queries or ACID transactions across multiple items (it only offers limited transactional APIs with restrictions). Option C is wrong because Amazon S3 is an object storage service with no support for ACID transactions, complex joins, or relational query capabilities. Option D is wrong because Amazon Redshift is a columnar data warehouse optimized for analytical queries on large datasets, not for transactional workloads requiring ACID compliance and complex joins at the row level.

557
MCQmedium

A data engineer is using Amazon Kinesis Data Streams to ingest real-time data. The stream has 4 shards and is receiving 2 MB/s of data. The engineer notices that the WriteProvisionedThroughputExceeded metric is increasing. The engineer wants to resolve this issue with minimal changes. What should the engineer do?

A.Increase the retention period of the stream to 168 hours.
B.Enable enhanced fan-out for the stream.
C.Increase the number of shards in the stream to 8.
D.Use a more random partition key to distribute data evenly across shards.
AnswerD

Correct. With 4 shards, the stream can handle up to 4 MB/s (1 MB/s per shard). Since the incoming rate is 2 MB/s, the throughput exceedance suggests that data is not evenly distributed, causing some shards to hit their 1 MB/s limit. Using a more random partition key, such as a UUID, distributes records evenly, resolving the issue without adding shards.

Why this answer

The WriteProvisionedThroughputExceeded metric indicates that some shards are exceeding their 1 MB/s write capacity. With 4 shards, total capacity is 4 MB/s, which is sufficient for 2 MB/s if evenly distributed. The issue is likely hot shards due to a non-uniform partition key.

Using a more random partition key, such as a UUID, ensures even distribution and resolves the throttling without adding shards.

Exam trap

The trap here is assuming that adding shards is always the first step, when the actual problem might be uneven data distribution across existing shards.

558
Multi-Selectmedium

A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format, and each record is approximately 5 KB. The company has set the buffer interval to 60 seconds and the buffer size to 5 MB. However, the data engineer observes that the delivery to S3 is delayed by up to 5 minutes during peak traffic. The engineer wants to reduce the delivery latency to under 1 minute. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Enable GZIP compression for the delivery stream.
B.Reduce the buffer size to 1 MB.
C.Increase the buffer size to 50 MB.
D.Convert the data format to Apache Parquet before delivery.
E.Reduce the buffer interval to 10 seconds.
AnswersB, E

Firehose flushes when either the buffer size or buffer interval is reached, whichever comes first. Lowering the size to 1 MB means the 5 MB threshold is hit sooner, triggering delivery earlier during peak traffic and cutting latency below one minute.

Why this answer

Option B is correct because Kinesis Data Firehose flushes data to S3 when either the buffer size OR the buffer interval is reached; lowering the buffer size from 5 MB to 1 MB makes the size threshold trigger much sooner, so smaller batches are delivered more frequently and latency drops. Option E is correct because reducing the buffer interval from 60 seconds to 10 seconds forces a flush at least every 10 seconds even if the buffer size is not reached, directly capping delivery latency well under 1 minute. Together, B and E ensure the delivery stream flushes on the shorter of the two thresholds, which is exactly what is needed to cut the observed 5-minute delays.

Option A does not help because GZIP compression only reduces payload size and can actually add CPU overhead, not reduce flush latency. Option C is wrong because increasing the buffer size to 50 MB would delay flushes further, worsening latency. Option D is wrong because converting to Parquet changes the storage format and query efficiency, not the Firehose buffering/flush timing that governs delivery latency.

Exam trap

The trap is assuming compression or format conversion (GZIP, Parquet) reduces delivery latency, when only the buffer size and buffer interval thresholds control when Firehose actually flushes data to S3.

559
MCQeasy

A data engineer applies the above IAM policy to a user. The user attempts to upload an object to the bucket 'my-data-lake' without specifying server-side encryption. What will happen?

A.The upload fails only if the bucket has a default encryption setting
B.The upload succeeds if the bucket policy allows unencrypted uploads
C.The upload succeeds because the policy allows s3:PutObject
D.The upload fails because the condition requires encryption
AnswerD

The policy's `s3:PutObject` statement carries a `StringNotEquals` condition on `s3:x-amz-server-side-encryption`, so a request omitting that header evaluates as non-compliant and is denied. Because the user supplied no encryption parameter, the condition fails and the upload is rejected with AccessDenied, satisfying the stem's unencrypted-upload constraint.

Why this answer

The IAM policy includes a condition that requires the `s3:x-amz-server-side-encryption` header to be present with a value of `AES256`. When the user attempts to upload an object without specifying server-side encryption, the condition is not satisfied, so the `s3:PutObject` permission is denied. This causes the upload to fail, regardless of any bucket default encryption settings or bucket policies.

Exam trap

The trap here is that candidates assume bucket default encryption automatically satisfies an IAM condition requiring the encryption header, but the condition checks the request headers, not the bucket's configuration.

How to eliminate wrong answers

Option A is wrong because the failure is due to the IAM policy condition, not the bucket's default encryption setting; even if the bucket has no default encryption, the IAM policy still denies the upload. Option B is wrong because the bucket policy is irrelevant here—the IAM policy explicitly denies unencrypted uploads via a condition, and bucket policies cannot override an explicit IAM deny. Option C is wrong because while the policy allows `s3:PutObject` in general, the condition `s3:x-amz-server-side-encryption` must be met; without it, the permission is effectively denied.

560
MCQhard

A data engineer is troubleshooting an Amazon Redshift cluster that has been experiencing slow query performance. The engineer checks the system tables and finds that many queries are waiting on 'wlm_queued' time. The cluster has 10 nodes and uses automatic WLM. What is the most likely cause?

A.Network bandwidth saturation between nodes.
B.Sorting operations are too expensive.
C.The number of concurrent queries exceeds the available query slots.
D.Insufficient disk space on the cluster.
AnswerC

High 'wlm_queued' wait time indicates queries are sitting in a workload management queue rather than executing. With automatic WLM, concurrency scaling is not configured here, so when simultaneous queries exceed the query slots automatic WLM allocates per queue, they queue. This directly matches the stem's observed wait event.

Why this answer

When queries show 'wlm_queued' time in Amazon Redshift, it indicates they are waiting in the Workload Management (WLM) queue before execution. With automatic WLM, the system dynamically manages concurrency, but if the number of concurrent queries exceeds the available query slots (determined by the cluster's memory and concurrency scaling settings), queries will be queued. This is the most direct cause of 'wlm_queued' wait events, as WLM queues queries when all slots are occupied.

Exam trap

The trap here is that candidates confuse 'wlm_queued' with resource contention (like CPU or I/O), but WLM queuing specifically indicates a concurrency limit, not a performance bottleneck during execution.

How to eliminate wrong answers

Option A is wrong because network bandwidth saturation between nodes would manifest as 'network' or 'distributed' wait events in system tables, not 'wlm_queued' time, which is purely a queuing delay. Option B is wrong because expensive sorting operations would appear as 'sort' or 'hash' wait times in query execution plans, not as WLM queue wait; sorting occurs after a query is assigned a slot. Option D is wrong because insufficient disk space would cause 'disk full' errors or 'resize' operations, not WLM queuing; Redshift's WLM queuing is independent of storage capacity.

561
MCQhard

A data engineer runs an AWS Glue job that reads from an Amazon Kinesis Data Stream and writes to Amazon S3. The job must process records in order per shard and must checkpoint progress so it can resume after a failure without reprocessing all records. Which Glue configuration BEST supports this?

A.Use a Glue streaming ETL job with a custom checkpoint table in Amazon DynamoDB and set `--starting_position` to `TRIM_HORIZON`.
B.Use a Glue streaming ETL job with `--starting_position` set to `LATEST` and disable job bookmarks to always read new records.
C.Use the Glue streaming ETL job type with `--starting_position` set to `TRIM_HORIZON` and rely on Glue's built-in checkpointing in the job bookmarks.
D.Use a standard Glue batch job with the Kinesis connector and set `--starting_position` to `LATEST` for each run.
AnswerC

Glue streaming ETL jobs support Kinesis Data Streams as a source and use job bookmarks to track processed records. Setting `--starting_position` to `TRIM_HORIZON` defines where to start on the first run, and bookmarks checkpoint progress so subsequent runs resume from the last processed record. This provides ordered, per-shard processing and fault-tolerant resumption without custom checkpoint code.

Why this answer

Glue streaming ETL jobs natively support Kinesis Data Streams and use job bookmarks to track progress per shard. Setting `--starting_position` to `TRIM_HORIZON` establishes the initial read point, and bookmarks checkpoint processed records so a failed job resumes without reprocessing everything. Batch jobs, disabled bookmarks, and custom checkpoint tables either lose data, reprocess, or add unnecessary overhead.

Exam trap

The trap here is confusing `--starting_position` on each run with true checkpointing, when resumption without reprocessing depends on job bookmarks being enabled.

562
MCQhard

A company runs a data warehouse on Amazon Redshift. The data engineer notices that some queries are running slowly. Upon reviewing the system tables, the engineer finds that the 'svv_table_info' shows high 'unsorted' percentage for several large tables. What is the MOST effective action to improve query performance?

A.Run the ANALYZE command on the tables.
B.Run the VACUUM command on the tables.
C.Change the distribution style of the tables to ALL.
D.Increase the number of nodes in the Redshift cluster.
AnswerB

Running VACUUM re-sorts rows into the table's defined sort key order, directly reducing the high unsorted percentage that svv_table_info reports. Because those large tables were loaded without sorted data, range-restricted scans read far more blocks than necessary; restoring sort order lets Redshift apply zone-map block pruning, cutting I/O for the slow queries.

Why this answer

A high 'unsorted' percentage in svv_table_info means many rows were added after the last sort and are not in sort-key order, forcing Redshift to scan more blocks and slowing range-restricted queries. VACUUM re-sorts the table and reclaims space, restoring the benefits of the sort key. This is the direct, most effective fix for the reported symptom.

Exam trap

DEA-C01 often tests the distinction between VACUUM (re-sorts and reclaims space) and ANALYZE (updates statistics), tricking candidates into choosing ANALYZE when the symptom is a high unsorted percentage.

How to eliminate wrong answers

Option A is wrong because ANALYZE only refreshes table statistics for the query planner; it does not re-sort rows or reduce the unsorted percentage. Option C is wrong because changing distribution style to ALL replicates the entire table to every node, which increases storage and load time and does not address unsorted data. Option D is wrong because adding nodes increases compute capacity but does not fix the underlying unsorted data that causes inefficient block scans.

563
MCQmedium

A data engineer is designing a data lake on Amazon S3 for a financial analytics workload. The raw data arrives as JSON files from an on-premises system. Analysts need to query the data using Amazon Athena with fast performance and minimal cost for queries that filter on a specific transaction date and customer ID. The engineer wants to convert the data to a columnar format that supports predicate pushdown and compression. Which storage format should the engineer choose?

A.Apache Avro
B.Apache Parquet
C.JSON
D.CSV
AnswerB

Apache Parquet is a columnar format that stores data by column, enabling Athena to read only the columns needed for a query. It supports predicate pushdown, so filters on transaction date and customer ID skip irrelevant data. Parquet also compresses well, reducing data scanned and query cost. This directly meets the requirement for fast, cost-effective queries on filtered columns.

Why this answer

Apache Parquet is a columnar storage format that allows Athena to read only the columns referenced in a query and to skip row groups using predicate pushdown. This reduces the amount of data scanned, lowering cost and improving performance for filters on transaction date and customer ID. Parquet also compresses efficiently, further reducing storage and scan volume.

Exam trap

The trap here is assuming that any compressed format will improve Athena performance, when only columnar formats like Parquet or ORC enable column pruning and predicate pushdown.

564
MCQmedium

A data engineer is configuring an AWS Glue ETL job that reads semi-structured JSON from Amazon S3 and must flatten nested arrays before loading into Amazon Redshift. During testing, the job fails with an AnalysisException stating that a column named 'events' cannot be resolved, even though the AWS Glue Data Catalog table shows the column. The engineer confirms the catalog table was created by a crawler and the S3 data is present. Which action will most directly resolve the schema resolution failure?

A.Call resolveChoice with the 'MatchCatalogSchema' option on the DynamicFrame before referencing the nested column.
B.Increase the number of DPUs allocated to the Glue job so the executor can read the full nested JSON payload.
C.Enable the 'Group files' option in the crawler and re-crawl the S3 prefix to consolidate the JSON into a single object.
D.Convert the source JSON to Parquet with an AWS Glue crawler and then point the job at the Parquet table.
AnswerA

When a crawler infers a schema that differs from what Spark infers at read time, nested fields can be dropped or renamed, producing unresolved-column errors. resolveChoice with MatchCatalogSchema forces the DynamicFrame to conform to the Data Catalog table definition, restoring the 'events' column so downstream transforms and the Redshift write can proceed.

Why this answer

The failure occurs because Spark's inferred schema at read time diverges from the Data Catalog definition created by the crawler, so nested fields such as 'events' are not resolvable in the DynamicFrame. Applying resolveChoice with MatchCatalogSchema reconciles the runtime schema with the catalog, restoring the column. Capacity changes and file grouping do not alter schema resolution, and crawlers do not perform format conversion.

Exam trap

The trap here is assuming any schema error means the catalog is wrong, when the real cause is a mismatch between Spark-inferred and catalog-inferred schemas that resolveChoice must reconcile.

565
MCQhard

A data engineer manages an AWS Lake Formation governed data lake. Analysts query tables through Amazon Athena and must see only rows where the region column equals their assigned region, while column-level restrictions must also hide a national ID column. The engineer grants table SELECT to the analysts' IAM role in Lake Formation. What should the engineer configure next?

A.Attach an IAM policy to the analysts' role that denies access to the national ID column and uses a condition on the region tag.
B.Create separate Athena views for each region and revoke direct table access, relying on view definitions for the row and column limits.
C.Enable Amazon Athena workgroup query result reuse and rely on Athena to mask the national ID column automatically.
D.Create a Lake Formation data filter that includes a row filter expression on region and excludes the national ID column, then grant SELECT with that data filter to the role.
AnswerD

Lake Formation data filters support both row-level filter expressions and column inclusion or exclusion in a single grant. Granting SELECT with the data filter applies the row predicate and hides the excluded column for that principal, which is exactly the granular control the scenario requires without duplicating tables.

Why this answer

Lake Formation data filters are the mechanism for combining row-level expressions with column inclusion or exclusion in a single SELECT grant. Granting the filter to the analysts' IAM role enforces both restrictions on the governed table, so analysts see only their region rows and never the national ID column, without duplicating tables or views.

Exam trap

The trap here is believing that IAM policies can enforce row or column restrictions on Lake Formation tables, when those controls live in Lake Formation grants.

566
Multi-Selectmedium

A data engineer is troubleshooting a slow-running Amazon Athena query. The query scans a large amount of data. Which TWO actions can improve query performance? (Choose TWO.)

Select 2 answers
A.Convert the data to Parquet or ORC format.
B.Enable encryption at rest.
C.Increase the Athena query timeout.
D.Partition the table on frequently filtered columns.
E.Use SELECT * to retrieve all columns.
AnswersA, D

Parquet and ORC are columnar formats, so Athena reads only the columns referenced by the query rather than every field in each row. This reduces the bytes scanned from S3, which is the dominant cost for the large scans described in the stem.

Why this answer

Option A is correct because converting data to columnar formats like Parquet or ORC lets Athena read only the columns referenced in the query and benefits from compression and predicate pushdown, drastically reducing the bytes scanned and thus improving performance and lowering cost. Option D is correct because partitioning the table on frequently filtered columns (e.g., date or region) enables partition pruning, so Athena skips scanning irrelevant partitions instead of reading the entire dataset. Option B is incorrect because encryption at rest protects stored data but does not reduce the amount of data scanned or speed up query execution.

Option C is incorrect because increasing the query timeout only allows a slow query to run longer; it does not make the query faster. Option E is incorrect because using SELECT * retrieves all columns, which increases the data scanned and worsens performance, especially with columnar formats.

567
Multi-Selectmedium

A company is designing a data ingestion pipeline for clickstream data from a website. The data must be ingested in near real-time. Which TWO services can be used together to build this pipeline?

Select 2 answers
A.Amazon Kinesis Data Streams
B.Amazon Simple Queue Service (SQS)
C.Amazon S3
D.Amazon Kinesis Data Firehose
E.Amazon DynamoDB
AnswersA, D

Amazon Kinesis Data Streams ingests clickstream events in near real time with low latency and durable retention. Paired with a consumer such as AWS Lambda or Kinesis Data Analytics, it forms the ingestion layer of a real-time pipeline, meeting the near-real-time requirement.

Why this answer

Amazon Kinesis Data Streams (A) is correct because it is designed for real-time streaming ingestion of high-volume data such as clickstream events, allowing producers to continuously write records that can be processed within milliseconds by consumers. Amazon Kinesis Data Firehose (D) is correct because it can consume streaming data and reliably load it into destinations like S3, Redshift, or Elasticsearch in near real-time, making it a natural complement to Kinesis Data Streams for building an end-to-end ingestion pipeline. Together, Kinesis Data Streams captures and buffers the clickstream events while Firehose delivers them to downstream storage or analytics services with minimal latency.

Amazon SQS (B) is a message queue for decoupling applications but is not purpose-built for real-time streaming analytics pipelines. Amazon S3 (C) is a storage service, not an ingestion mechanism, and Amazon DynamoDB (E) is a NoSQL database, so neither serves as the streaming ingestion component required here.

568
MCQhard

A company needs to share a dataset stored in an S3 bucket with a partner account. The dataset contains sensitive information, so the company wants to ensure that the partner account can only access the data using a specific VPC endpoint in the partner's account. Which S3 bucket policy condition key should be used?

A.aws:SourceVpc
B.aws:SourceArn
C.aws:SourceIp
D.aws:SourceVpce
AnswerD

The `aws:SourceVpce` condition key restricts access to requests arriving through a named VPC endpoint, satisfying the requirement that the partner reach the bucket only via their specific endpoint. Because it evaluates the endpoint ID itself, requests from any other endpoint or the public internet are denied, even if the partner's IAM permissions allow S3 access.

Why this answer

The aws:SourceVpce condition key restricts access to a specific VPC endpoint, ensuring the partner can only access the data through that endpoint. Option A is wrong because aws:SourceVpc restricts to a VPC, not a VPC endpoint. Option B is wrong because aws:SourceArn restricts to a resource ARN, not a network endpoint.

Option C is wrong because aws:SourceIp restricts to an IP address, which does not meet the requirement.

569
MCQhard

A data engineer manages an Amazon Redshift cluster that experiences performance degradation during peak hours due to concurrent long-running queries and short ad-hoc queries competing for resources. The engineer wants to isolate the workloads so that short queries are not blocked by long-running ones, and to ensure that each workload gets a guaranteed share of memory and CPU. Which Redshift feature should the engineer implement?

A.Redshift Spectrum
B.Workload Management (WLM) with query queues
C.Short Query Acceleration (SQA)
D.Concurrency Scaling
AnswerB

WLM allows defining multiple query queues with different memory and concurrency settings. By configuring separate queues for long-running and short ad-hoc queries, the engineer can isolate workloads, ensuring that short queries are not blocked by long ones and that each queue gets a guaranteed share of resources. This directly addresses the scenario.

Why this answer

Workload Management (WLM) enables the creation of separate query queues with configurable memory and concurrency, allowing isolation of long-running and short ad-hoc queries. This ensures that short queries are not blocked and that each workload receives a guaranteed share of resources. Other features like Concurrency Scaling, SQA, and Redshift Spectrum do not provide the same level of workload isolation.

Exam trap

The trap here is confusing Concurrency Scaling or SQA with workload isolation; these features improve performance for certain query types but do not provide dedicated resource queues.

570
MCQhard

A data engineer is troubleshooting a slow Amazon Redshift query. The query scans a large table with interleaved sort keys. The engineer notices that the query plan shows a sequential scan instead of a range-restricted scan. What is the MOST likely reason?

A.The table has not been vacuumed and reindexed after large data loads.
B.The table has a poor distribution key (DISTKEY) causing data skew.
C.The table uses compression encodings that prevent range-restricted scans.
D.The workload management (WLM) queue is configured with too few query slots.
AnswerA

Interleaved sort keys rely on accurate zone maps; after large loads, unsorted or deleted rows degrade them, so Redshift cannot skip blocks and falls back to a sequential scan. Vacuum and reindex restore the range-restricted scan.

Why this answer

Interleaved sort keys in Amazon Redshift rely on zone maps that must be refreshed via VACUUM REINDEX after significant data changes. Without reindexing, the zone maps become stale and the query planner cannot skip blocks, forcing a sequential scan instead of a range-restricted scan. Running VACUUM REINDEX rebuilds the interleaved sort metadata so range predicates can prune blocks effectively.

Exam trap

The trap is blaming distribution or compression for scan-type problems; the exam tests whether you know that interleaved sort keys specifically require VACUUM REINDEX to keep zone maps usable.

How to eliminate wrong answers

Option B is wrong because a poor DISTKEY causes data skew and uneven node utilization, but it does not prevent range-restricted scans — that is a sort-key/zone-map concern. Option C is wrong because compression encodings (e.g., AZ64, ZSTD) do not block range-restricted scans; Redshift can filter on compressed blocks using zone maps. Option D is wrong because WLM queue slot configuration affects query concurrency and queuing, not the planner's ability to perform range-restricted scans.

571
MCQmedium

Refer to the exhibit. An AWS Glue ETL job is failing with an OutOfMemoryError. The job reads from Amazon S3 and performs a GROUP BY on a large dataset. Which change should the data engineer make to resolve this error?

A.Use coalesce to reduce the number of partitions.
B.Increase the number of DPUs allocated to the Glue job.
C.Increase the number of partitions in the DataFrame.
D.Use repartition to increase the number of partitions.
AnswerB

OutOfMemoryError during a GROUP BY on a large dataset occurs because each executor has insufficient memory for the shuffle and aggregation. Increasing DPUs adds more executors and memory, distributing the GROUP BY workload and preventing the driver or executor from exhausting heap during the shuffle phase.

Why this answer

The OutOfMemoryError in an AWS Glue ETL job performing a GROUP BY on a large dataset indicates that the executors do not have enough memory to handle the shuffle operations required for aggregation. Increasing the number of DPUs (Data Processing Units) allocated to the Glue job increases the total memory and compute resources available, allowing the job to process larger partitions without running out of memory.

Exam trap

The trap here is that candidates often confuse partition tuning (coalesce/repartition) with resource allocation, mistakenly thinking that adjusting partitions alone can fix memory errors without increasing the underlying compute and memory capacity.

How to eliminate wrong answers

Option A is wrong because using coalesce to reduce the number of partitions would decrease parallelism and concentrate data into fewer partitions, potentially worsening memory pressure and making the OutOfMemoryError more likely. Option C is wrong because increasing the number of partitions in the DataFrame without adding more resources (DPUs) would spread data across more tasks but still rely on the same total memory, which does not resolve the underlying memory shortage. Option D is wrong because repartitioning to increase the number of partitions similarly does not add memory; it only redistributes data, which can even increase shuffle overhead and exacerbate memory issues.

572
Multi-Selectmedium

A company uses S3 to store sensitive data. Which TWO S3 features can be used to protect data at rest?

Select 2 answers
A.S3 Versioning
B.Server-Side Encryption with S3 Managed Keys (SSE-S3)
C.Server-Side Encryption with AWS KMS (SSE-KMS)
D.S3 Transfer Acceleration
E.S3 Object Lock
AnswersB, C

SSE-S3 encrypts objects at rest using AES-256, with Amazon S3 managing the key material and rotation automatically. This satisfies the stem's data-at-rest protection requirement without the company provisioning or maintaining keys in AWS KMS, removing key-management overhead entirely.

Why this answer

SSE-S3 (option B) is correct because it encrypts objects at rest using AES-256 with keys fully managed by Amazon S3, directly protecting stored data without any customer key management. SSE-KMS (option C) is also correct because it encrypts objects at rest using AWS KMS customer master keys, adding key control, audit trails via CloudTrail, and separation of duties. S3 Versioning (A) only keeps multiple variants of an object to aid recovery; it does not encrypt data.

S3 Transfer Acceleration (D) speeds up uploads/downloads over AWS edge locations and is a transfer optimization, not encryption at rest. S3 Object Lock (E) enforces WORM retention to prevent deletion or modification, which is immutability, not encryption.

573
MCQmedium

A data engineer needs to load data from an Amazon DynamoDB table into an Amazon S3 data lake nightly. The table is large and the engineer wants to avoid consuming provisioned read capacity on the live table. Which approach should the engineer use?

A.Create an Amazon ElastiCache for Redis cluster and replicate the DynamoDB table into it, then export from Redis to Amazon S3.
B.Use the DynamoDB export to Amazon S3 feature to export a point-in-time snapshot to S3 without consuming read capacity.
C.Enable DynamoDB Streams on the table and use an AWS Lambda function to write each change to Amazon S3 in near real time.
D.Use AWS Glue with a DynamoDB connection and read the table directly with the DynamoDB reader, accepting the consumed read capacity.
AnswerB

DynamoDB export to S3 creates a full point-in-time snapshot in S3 without consuming provisioned read capacity on the table, because it reads from the underlying storage. It is designed for analytics and backup use cases, and the resulting data can be queried or transformed directly in S3.

Why this answer

DynamoDB export to S3 produces a point-in-time snapshot directly in Amazon S3 without consuming read capacity on the source table, making it ideal for nightly analytics loads that must not impact production traffic. The exported data is written in a format suitable for downstream processing by services such as Athena or Glue.

Exam trap

The trap here is assuming that any read of DynamoDB consumes read capacity, when the export-to-S3 feature reads from storage and bypasses provisioned throughput.

574
MCQhard

A data engineer is using AWS Glue to read from an Amazon S3 bucket that contains data in Apache Parquet format, partitioned by year/month/day. The Glue job needs to read only the data for the last 7 days. The engineer wants to minimize the amount of data scanned and improve job performance. Which approach should be used to filter the partitions efficiently?

A.Create a separate Glue Data Catalog table for each day's data and read from the last 7 tables.
B.Read the entire dataset into a DynamicFrame and then apply a filter transformation on the partition columns.
C.Use a pushdown predicate in the Glue job's 'create_dynamic_frame.from_catalog' call to filter on partition columns.
D.Use the 'glueContext.read_from_options' with a 'filter' parameter to specify the partition range.
AnswerC

This option is correct because AWS Glue supports pushdown predicates when reading from the AWS Glue Data Catalog. By specifying a filter expression on partition columns (e.g., 'year >= 2023 AND month = 10 AND day BETWEEN 1 AND 7'), Glue pushes the filter down to the data source, so only the relevant partitions are read from S3. This significantly reduces the amount of data scanned and improves performance.

Why this answer

Using a pushdown predicate in the 'create_dynamic_frame.from_catalog' call allows AWS Glue to filter partitions at the source, so only the relevant S3 partitions for the last 7 days are read. This minimizes data scanned and improves job performance by avoiding loading unnecessary data.

Exam trap

The trap here is assuming that filtering after reading the data is equivalent to filtering at the source, when in fact pushdown predicates are required to avoid scanning all partitions.

575
MCQmedium

A data engineer needs to migrate an on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which approach should they use?

A.Use mysqldump to export the database and import into RDS.
B.Use AWS Database Migration Service (DMS) with ongoing replication from the source database.
C.Create an RDS read replica and promote it.
D.Use AWS Schema Conversion Tool (SCT) to convert the schema and then copy data.
AnswerB

DMS performs a full load then applies ongoing change data capture replication from the MySQL source, keeping the target synchronised until cutover. This continuous replication is what achieves minimal downtime rather than a one-off snapshot migration.

Why this answer

AWS DMS with ongoing replication (change data capture, CDC) is the correct approach because it allows continuous synchronization from the on-premises MySQL source to the RDS target, enabling a cutover with minimal downtime. Unlike one-time export/import tools, DMS captures ongoing changes during the migration, so the target stays up-to-date until you switch over.

Exam trap

The trap here is that candidates confuse 'minimal downtime' with 'zero data loss' and assume a simple dump/import or a read replica (which only works for RDS-to-RDS) is sufficient, overlooking the need for ongoing replication to keep the target synchronized during the migration window.

How to eliminate wrong answers

Option A is wrong because mysqldump performs a logical backup that requires the source database to be locked or read-only during the dump, causing significant downtime; it also does not support ongoing replication. Option C is wrong because RDS read replicas can only be created from an existing RDS instance, not from an on-premises database, and promoting a replica does not migrate data from an external source. Option D is wrong because AWS Schema Conversion Tool (SCT) is designed for heterogeneous migrations (e.g., Oracle to Aurora) and does not handle data replication; for a homogeneous MySQL-to-MySQL migration, SCT is unnecessary and does not provide ongoing sync.

576
MCQhard

A company runs an e-commerce platform that generates clickstream data from user interactions on their website. The data is sent as JSON objects via HTTP POST to an API Gateway endpoint, which triggers a Lambda function that writes each record to a Kinesis Data Stream (100 shards). A second Lambda function consumes the stream, transforms the data (enriches with geolocation from a DynamoDB table), and writes to a Kinesis Data Firehose delivery stream that delivers Parquet files to an S3 data lake every 5 minutes. The system has been working for months, but recently the Firehose delivery stream started showing 'DeliveryFailed' errors for a subset of records. The errors point to 'InvalidData' from the Lambda transformation. The engineer reviews the Lambda transformation code and notices that the geolocation lookup occasionally fails because the DynamoDB table has a throttling issue. The engineer needs to handle these failures gracefully so that records that fail enrichment are still delivered to S3 with a null geolocation field, without blocking other records. Which course of action should the engineer take?

A.Configure the Kinesis Data Firehose delivery stream to send failed records to a dead-letter queue (DLQ) for later reprocessing.
B.Modify the Lambda function to send failed records to a separate Kinesis Data Stream for manual processing.
C.Modify the Lambda function to catch exceptions during the geolocation lookup, set the geolocation field to null, and continue processing the record.
D.Increase the read capacity units (RCUs) on the DynamoDB table to eliminate throttling.
AnswerC

Wrapping the DynamoDB geolocation lookup in try/catch lets the function substitute a null geolocation and emit the record rather than throwing, which is what currently produces the InvalidData delivery failures. This satisfies the requirement that unenriched records still reach S3 without blocking other records in the batch.

Why this answer

The requirement is to keep records flowing to S3 with a null geolocation when enrichment fails, without blocking other records. Wrapping the DynamoDB lookup in a try/catch inside the Lambda transformation, setting geolocation to null on failure, and returning the record as 'Ok' ensures Firehose treats it as successfully transformed and delivers it. This is the standard graceful-degradation pattern for Firehose Lambda transformations.

Exam trap

DEA-C01 often tests whether candidates choose 'route failures elsewhere' (DLQ, separate stream) versus 'handle failures inline and continue' — the trap is picking DLQ because it sounds robust, when the requirement is to still deliver the record with a null field.

How to eliminate wrong answers

Option A is wrong because Firehose's DLQ/S3 error bucket captures records that fail transformation or delivery, but the requirement is to still deliver those records to S3 with a null geolocation — sending them to a DLQ means they are not delivered to the data lake as intended and require separate reprocessing. Option B is wrong because routing failed records to a separate Kinesis Data Stream adds operational complexity and still does not deliver them to S3 with a null field; it also does not prevent the Firehose DeliveryFailed errors. Option D is wrong because increasing DynamoDB RCUs addresses the root cause of throttling but does not handle the failure gracefully — throttling can still occur (e.g., during bursts or hot partitions), and the requirement is explicitly about handling failures so records are not lost or blocked.

577
MCQhard

A company uses DynamoDB with global tables in two AWS Regions. The data engineer observes that a write to the table in us-east-1 is not immediately visible in a read from eu-west-1. What is the most likely reason?

A.Replication between regions is eventually consistent.
B.The read is using strongly consistent reads.
C.There is a write conflict that needs to be resolved.
D.DynamoDB Streams is not enabled on the table.
AnswerA

DynamoDB global tables replicate asynchronously, so a write committed in us-east-1 propagates to eu-west-1 only after a short delay. Reads in the remote Region can therefore return stale data until replication completes, which explains the observed lag.

Why this answer

DynamoDB global tables use asynchronous replication between regions. When a write occurs in us-east-1, the change is propagated to eu-west-1 with a replication lag that is typically sub-second but not instantaneous. Reads in eu-west-1 are eventually consistent by default, meaning they may not reflect the most recent write until replication completes.

This is the expected behavior of DynamoDB global tables, which prioritize availability and partition tolerance over immediate consistency across regions.

Exam trap

The trap here is that candidates often assume DynamoDB global tables provide strong consistency across regions because they are familiar with single-region strongly consistent reads, but the exam tests the specific knowledge that cross-region replication is always eventually consistent and that strongly consistent reads are only valid within the same region.

How to eliminate wrong answers

Option B is wrong because strongly consistent reads would actually increase the chance of seeing stale data in a cross-region scenario, as they are only guaranteed to return the most recent write within the same region, not across regions; DynamoDB does not support cross-region strongly consistent reads. Option C is wrong because write conflicts in global tables are automatically resolved using a last-writer-wins algorithm based on the timestamp, and they do not cause writes to be invisible; a conflict would result in one write being overwritten, not a delay in visibility. Option D is wrong because DynamoDB Streams is not required for global table replication; global tables use their own internal replication mechanism, and enabling Streams is optional for change data capture or triggering Lambda functions, not for the core replication functionality.

578
MCQhard

A data engineer is using AWS Glue job bookmarks to process incremental data from Amazon S3. The job reads from a partitioned S3 path and writes to Amazon Redshift. After a recent run, the engineer notices that some new partitions were not processed. The job bookmark state shows that the job has already processed up to a certain timestamp. What is the most likely reason for the missing partitions?

A.The new partitions were created with a timestamp earlier than the last processed timestamp.
B.The Glue job bookmark was not enabled for the S3 data source.
C.The Redshift cluster's copy command does not support partitioned data.
D.The Glue job's IAM role lacks permissions to read the new partitions.
AnswerA

Glue job bookmarks track the last processed timestamp or partition. If new data arrives with an older timestamp (e.g., late-arriving data), it may be considered already processed and skipped. This is a common issue with bookmarks when data is not strictly ordered by time. The engineer should adjust the bookmark logic or handle late data.

Why this answer

Glue job bookmarks use timestamps or partition values to determine what data has already been processed. If new partitions have timestamps older than the last processed timestamp, they are considered already processed and skipped. This often happens with late-arriving data.

The engineer should either adjust the bookmark to use partition-based tracking or handle late data separately.

Exam trap

The trap here is assuming that all new data will be processed regardless of timestamp, but bookmarks rely on ordering and can miss late-arriving partitions.

579
MCQmedium

Refer to the exhibit. A data engineer is using a Kinesis Data Stream with one shard. The application writes 2000 records per second, each 1 KB. The put record calls are frequently throttled. What is the most likely cause?

A.The stream has only one shard, which limits writes to 1000 records per second
B.The retention period of 24 hours is too short
C.The stream uses KMS encryption, causing additional latency
D.Enhanced monitoring is not enabled, causing performance issues
AnswerA

A single shard caps writes at 1,000 records per second and 1 MB per second. The application attempts 2,000 records per second (about 2 MB per second), exceeding both limits, so Kinesis throttles the put calls. Adding shards raises the ingest capacity to match the required throughput.

Why this answer

A Kinesis Data Stream shard has a write throughput limit of 1,000 records per second (or 1 MB per second). Since the application is writing 2,000 records per second (each 1 KB) to a single shard, it exceeds the shard's record-per-second quota, causing the PutRecord calls to be throttled. The solution is to increase the number of shards to at least two to distribute the load.

Exam trap

The DEA-C01 exam often tests the misconception that throttling is caused by encryption latency or monitoring settings, but the real trap is forgetting that each shard has a hard limit of 1,000 records per second, regardless of other configurations.

How to eliminate wrong answers

Option B is wrong because the retention period (default 24 hours, max 365 days) controls how long records are stored, not the write throughput; throttling is unrelated to retention. Option C is wrong because KMS encryption adds latency to encrypt/decrypt operations but does not reduce the shard-level write limit of 1,000 records per second; throttling is a capacity issue, not a latency issue. Option D is wrong because enhanced monitoring provides detailed metrics (e.g., user, request, stream-level) but does not affect the shard's write throughput limits; throttling occurs regardless of monitoring settings.

580
MCQmedium

A data engineer is using AWS Glue to run a PySpark ETL job that processes millions of small JSON files stored in an Amazon S3 bucket. The job is experiencing high memory usage and failing with 'Container killed by YARN for exceeding memory limits'. The engineer wants to optimize the job to handle the data more efficiently without changing the source data format. Which solution will MOST effectively reduce memory usage and improve performance?

A.Convert the JSON files to Parquet format using an AWS Glue crawler before running the ETL job.
B.Use AWS Glue's groupFiles and groupSize options to combine small files into larger chunks during the read operation.
C.Increase the number of DPUs allocated to the AWS Glue job to provide more memory per executor.
D.Enable AWS Glue job bookmarks to track processed files and avoid reprocessing.
AnswerB

The groupFiles and groupSize options in AWS Glue allow the job to coalesce multiple small files into larger groups, reducing the number of tasks and memory overhead per file. This directly addresses the inefficiency of processing many small files by minimizing the number of Spark partitions and the associated memory pressure, leading to better performance and lower resource consumption.

Why this answer

The groupFiles and groupSize options in AWS Glue are specifically designed to handle large numbers of small files by grouping them into larger chunks during read, which reduces the number of Spark partitions and memory overhead. This optimization directly targets the root cause of the memory failures and improves job efficiency without altering the source data format.

Exam trap

The trap here is assuming that increasing DPUs or enabling bookmarks will solve the memory issue, when the real problem is the inefficiency of processing many small files.

581
Multi-Selecthard

A company uses Amazon Redshift for analytics. They notice that some queries are slow due to data redistribution. The data engineer wants to minimize data movement across nodes. Which table design strategy should be used? (Choose TWO.)

Select 2 answers
A.Set the distribution style to AUTO for all tables.
B.Define compound sort keys on frequently filtered columns.
C.Choose a distribution key that matches the join key for large tables.
D.Use EVEN distribution for all tables.
E.Use distribution style ALL for small dimension tables.
AnswersC, E

When two large tables share the same distribution key as their join column, matching rows are already co-located on the same slice. Redshift avoids the broadcast or shuffle step during joins, eliminating the cross-node data redistribution that was slowing queries.

Why this answer

Option C is correct because choosing a distribution key that matches the join key on large tables colocates matching rows on the same compute node slice, so joins between those tables can be performed locally without broadcasting or redistributing data across nodes. Option E is correct because using distribution style ALL replicates small dimension tables to every node, eliminating the need to redistribute the dimension during joins with large fact tables and thereby minimizing cross-node data movement. Option A is not ideal here because AUTO lets Redshift decide and may still choose EVEN or KEY distribution that results in redistribution for some workloads, rather than guaranteeing join-key alignment.

Option B addresses sort keys, which optimize range-restricted scans and merge joins but do not control how rows are distributed across nodes, so they do not directly reduce redistribution. Option D is incorrect because EVEN distribution spreads rows round-robin regardless of join keys, which typically forces data redistribution during joins and can worsen the problem.

Exam trap

The trap here is that candidates often confuse distribution keys with sort keys, thinking that sorting alone can reduce data movement, or they assume AUTO distribution always optimizes for joins, when in fact it may default to EVEN or ALL without guaranteeing collocation for specific join patterns.

582
Drag & Dropmedium

Order the steps to set up an Amazon EMR cluster for processing data in S3 using Spark.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, prepare the S3 bucket. Then launch the EMR cluster with Spark, configure instances, submit the job, and terminate.

583
MCQmedium

A company uses AWS Lambda to process messages from an Amazon SQS queue. The messages contain JSON payloads that need to be transformed and written to an Amazon DynamoDB table. Recently, the Lambda function has been timing out and messages are being sent to the dead-letter queue (DLQ). What is the BEST way to troubleshoot and resolve this issue?

A.Use a standard SQS queue instead of a DLQ to reprocess failed messages automatically.
B.Increase the Lambda function timeout and monitor DynamoDB write capacity to ensure it is not throttling.
C.Switch the SQS queue to a FIFO queue to ensure exactly-once processing.
D.Increase the visibility timeout of the SQS queue to 30 minutes.
AnswerB

Timeouts and DLQ routing usually stem from the function exceeding its configured duration or downstream throttling. Raising the timeout addresses slow execution, while monitoring DynamoDB write capacity reveals throttling that silently extends processing time, letting you right-size capacity or batch writes.

Why this answer

The Lambda function is timing out, which suggests the function's execution duration is exceeding its configured timeout. Increasing the timeout gives the function more time to process messages. Additionally, if DynamoDB write capacity is insufficient, throttling can cause retries that further delay processing, so monitoring and possibly increasing write capacity units addresses the root cause of timeouts.

Exam trap

The trap here is that candidates often focus on SQS queue configuration (visibility timeout, queue type) rather than addressing the actual performance bottleneck in the Lambda function or downstream DynamoDB service.

How to eliminate wrong answers

Option A is wrong because using a standard SQS queue instead of a DLQ does not automatically reprocess failed messages; it only changes the queue type and does not address the underlying timeout or throttling issue. Option C is wrong because switching to a FIFO queue enforces exactly-once processing and message ordering, but it does not resolve timeouts or DynamoDB throttling; it could even introduce additional latency. Option D is wrong because increasing the visibility timeout to 30 minutes only delays when a message becomes visible again after a failure, but it does not fix the root cause of timeouts or throttling; it may mask the problem.

584
MCQhard

A data engineer is troubleshooting an AWS Glue job that writes data to an Amazon S3 bucket in Parquet format. The job runs successfully but the output files are smaller than the configured 'groupFiles' size. The engineer has set 'groupFiles' to 'inPartition' and 'groupSize' to 1 GB. The input data is 10 GB in a single partition. What is the most likely reason for the small files?

A.The 'groupFiles' parameter is deprecated in the current Glue version.
B.The 'groupFiles' parameter only affects the input read phase, not the output write phase.
C.The 'groupFiles' parameter is misspelled or set incorrectly.
D.The engineer must also set 'repartition' to 1 to merge output files.
AnswerB

Grouping coalesces small input files during reading but does not control output file size.

Why this answer

'groupFiles' only works when the input data is already small and needs to be coalesced. However, if the input is large and the job writes output, the output file size is determined by the number of Spark partitions, not grouping. The grouping feature only applies to reading input files.

Option A is wrong because the setting is correct. Option C is wrong because grouping is a read-time feature, not write-time. Option D is wrong because grouping does not require repartitioning.

585
MCQmedium

A company is migrating its on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is 2 TB and runs on a single server. Which AWS service should be used for the migration?

A.AWS DataSync
B.Amazon S3 Transfer Acceleration
C.AWS Database Migration Service (DMS)
D.AWS Snowball Edge
AnswerC

DMS provides minimal downtime migration with change data capture.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it supports homogeneous migrations from Oracle to Amazon Aurora PostgreSQL with minimal downtime using ongoing replication (change data capture). DMS can handle a 2 TB source database by performing a full load followed by continuous replication of changes from the Oracle redo logs, allowing the target Aurora database to stay nearly in sync until cutover.

Exam trap

The trap here is that candidates may confuse AWS DataSync or Snowball Edge as viable for database migrations because they handle large data volumes, but they lack the schema conversion and ongoing replication capabilities required for minimal-downtime database migrations.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is designed for moving large datasets over the network between on-premises storage and AWS storage services (e.g., S3, EFS, FSx), not for database migrations with schema conversion and ongoing replication. Option B is wrong because Amazon S3 Transfer Acceleration only speeds up uploads to S3 buckets over the internet using optimized network paths; it does not migrate databases or handle schema conversion and CDC. Option D is wrong because AWS Snowball Edge is a physical data transfer device for offline bulk data movement, which would introduce significant downtime and cannot perform live replication or schema conversion for a database migration.

586
MCQhard

A data engineer manages an AWS Glue Data Catalog shared across teams. Analysts in one team must be able to query only the sales database and its tables, while another team owns the marketing database. The engineer wants permissions managed centrally in Lake Formation and wants the analysts to be able to create their own temporary tables but not alter the sales tables. Which combination of Lake Formation grants should the engineer apply?

A.Grant ALL on the sales database to the analysts and rely on IAM policies to restrict which tables they can alter.
B.Grant SELECT on the sales tables and DESCRIBE on the sales database, and grant ALL on the marketing database to the analysts.
C.Grant DESCRIBE and SELECT on the sales database and tables, and grant ALTER on the sales database so analysts can create temporary tables there.
D.Grant DESCRIBE and SELECT on the sales database and its tables to the analysts, and grant CREATE_TABLE on a separate scratch database for their temporary tables.
AnswerD

Granting DESCRIBE and SELECT on the sales database and tables lets analysts discover and query that data without altering it, since ALTER and DROP are separate permissions not granted. CREATE_TABLE on a dedicated scratch database allows them to build temporary tables without touching the sales database. This least-privilege combination satisfies both query access and safe self-service while keeping marketing data inaccessible.

Why this answer

Least privilege in Lake Formation separates read permissions from modification and creation rights. DESCRIBE and SELECT on the sales database and tables allow discovery and querying, while withholding ALTER and DROP prevents changes to production tables. CREATE_TABLE on a dedicated scratch database lets analysts build temporary tables safely, keeping marketing data and the sales schema protected and centrally governed.

Exam trap

The trap here is granting ALL or ALTER on a database to enable temporary tables, which also hands analysts the ability to modify or drop production tables.

587
Multi-Selecteasy

A company uses AWS Glue to run ETL jobs daily. The data engineer wants to reduce costs by optimizing the job configuration. Which two actions will help reduce costs? (Choose TWO.)

Select 2 answers
A.Use G.1X worker type instead of G.2X
B.Increase the job timeout to 48 hours
C.Enable Spark UI logging for debugging
D.Reduce the number of DPUs allocated to the job if the data volume is small
E.Increase the number of job retries to handle transient failures
AnswersA, D

G.1X is half the cost of G.2X.

Why this answer

G.1X workers provide 16 GB of memory and 4 vCPUs, while G.2X workers provide 32 GB and 8 vCPUs. For many ETL jobs, especially those with smaller data volumes or less complex transformations, G.1X workers are sufficient, and using them instead of G.2X directly reduces the cost per DPU-hour since AWS Glue pricing is based on DPU capacity. This optimization lowers costs without sacrificing performance if the job is not memory- or CPU-bound.

Exam trap

The trap here is that candidates often confuse cost optimization with reliability improvements, such as retries or timeouts, and may overlook that reducing worker size or DPU count directly lowers resource consumption and cost.

588
MCQmedium

A data engineer sees the CloudWatch log entry in the exhibit for a Lambda function that processes data from an Amazon SQS queue. What is the MOST likely cause of the timeout?

A.The Lambda function's reserved concurrency is set too low.
B.The Lambda function is running out of memory.
C.The Lambda function's timeout is too short for the processing required.
D.The SQS queue's visibility timeout is set too low.
AnswerC

Lambda terminates execution once the configured timeout elapses, so a function whose processing exceeds that limit fails mid-batch. The SQS-triggered invocation needs enough duration for the workload; extending the timeout directly addresses the premature termination shown in the log.

Why this answer

The CloudWatch log entry shows a Lambda invocation timing out, which directly indicates the function's configured timeout is shorter than the time needed to process the SQS message. Lambda timeouts are a hard limit (max 15 minutes), and when exceeded, the invocation is terminated and the message returns to the queue. Increasing the timeout (or optimizing the code) resolves the issue.

Exam trap

DEA-C01 often tests the confusion between Lambda timeout, SQS visibility timeout, and concurrency throttling — candidates must read the log message carefully to distinguish a timeout from a throttling or memory error.

How to eliminate wrong answers

Option A is wrong because low reserved concurrency causes throttling (429 errors) and messages to remain in the queue, not invocation timeouts — the log would show throttling, not a timeout. Option B is wrong because running out of memory produces an 'OutOfMemoryError' or 'Runtime exited with error: signal: killed' message, not a timeout. Option D is wrong because a low SQS visibility timeout causes duplicate processing (messages reappear before processing completes), not a Lambda timeout — the Lambda would still complete its invocation.

589
MCQmedium

A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an on-premises Oracle database to an Amazon S3 bucket in near real time. The source table has a primary key and the database is configured for ARCHIVELOG mode. The engineer needs change data capture (CDC) to capture INSERT, UPDATE, and DELETE operations. Which AWS DMS task setting should be used to capture ongoing changes?

A.Set the migration type to 'full-load' and configure ongoing replication using Oracle GoldenGate.
B.Set the migration type to 'cdc' only and rely on the Oracle redo logs without supplemental logging.
C.Set the migration type to 'full-load-and-cdc' and enable supplemental logging on the source table.
D.Set the migration type to 'full-load-and-cdc' and increase the commit rate to reduce latency.
AnswerC

For Oracle CDC, AWS DMS requires supplemental logging to be enabled at the database or table level to capture column data for updates and deletes. The migration type 'full-load-and-cdc' performs an initial full load and then continues capturing ongoing changes, satisfying the requirement for near real-time replication of INSERT, UPDATE, and DELETE operations.

Why this answer

To capture ongoing changes from Oracle to S3 with AWS DMS, the task must be configured with migration type 'full-load-and-cdc' and supplemental logging must be enabled on the source. Supplemental logging provides the additional column data needed for UPDATE and DELETE operations. Without it, CDC cannot function correctly.

The other options either use incorrect migration types or omit the critical supplemental logging requirement.

Exam trap

The trap here is assuming that setting the migration type to 'cdc' alone is sufficient for Oracle, while overlooking the mandatory supplemental logging requirement on the source database.

590
MCQeasy

A data engineer wants to ensure that only users with a specific tag (e.g., "Department": "DataEngineering") can access an S3 bucket. How can this be enforced?

A.Use a bucket policy with aws:PrincipalTag condition
B.Use S3 object tags and a bucket policy condition
C.Attach an IAM policy to each user with the tag
D.Use S3 Object Lambda to check user tags
AnswerA

A bucket policy with the `aws:PrincipalTag` condition key evaluates the tag attached to the calling principal's identity, so only users tagged `Department: DataEngineering` satisfy the condition and gain access. This directly enforces the stem's tag-based restriction at the resource, without relying on group membership or role naming.

Why this answer

S3 bucket policies support condition keys like aws:PrincipalTag, which allow access control based on tags attached to IAM principals (users or roles). Option A is correct because it uses aws:PrincipalTag in the bucket policy to restrict access to users with the specific tag. Option B is incorrect because S3 object tags are for objects, not principals, and cannot be used to filter users.

Option C is incorrect because attaching an IAM policy to each user is less scalable and does not leverage the bucket policy's centralized control. Option D is incorrect because S3 Object Lambda is for modifying data during retrieval, not for access control decisions.

Exam trap

Be careful not to confuse principal tags with resource tags. The condition 'aws:PrincipalTag' checks the requester's IAM user/role tags, while 's3:ExistingObjectTag' checks tags on the S3 object itself. This question tests the distinction.

591
MCQeasy

A data engineer is setting up an Amazon Redshift cluster and needs to load data from Amazon S3. The data is in CSV format and contains a large number of rows. The engineer wants to achieve the fastest possible load time. Which method should the engineer use?

A.Use Amazon Kinesis Data Firehose to stream the data from S3 to Redshift.
B.Use AWS Database Migration Service (AWS DMS) to migrate the data from S3 to Redshift.
C.Use the COPY command to load the data from Amazon S3 into the Redshift cluster.
D.Use the INSERT INTO command to load the data row by row from S3.
AnswerC

The COPY command is the most efficient way to load large datasets from Amazon S3 into Redshift. It leverages massively parallel processing, automatically distributes the load across all nodes, and can parse CSV files directly. It also supports compression and columnar formats for even faster loads, making it ideal for this scenario.

Why this answer

The COPY command is specifically designed for bulk loading data from Amazon S3 into Redshift. It uses parallel processing to load data quickly and efficiently, supports CSV format, and can handle large volumes. Other methods like DMS or Firehose are not optimized for this use case, and INSERT is too slow.

Exam trap

The trap here is assuming that any AWS data transfer service can load data quickly, but only COPY is optimized for bulk loading into Redshift from S3.

592
MCQhard

Refer to the exhibit. A data engineer runs the command on an object in S3. The engineer expected the object to have a tag 'type=raw' but sees no metadata. What is the likely cause?

A.Object tags are not returned by head-object; use get-object-tagging instead
B.The S3 bucket is in a different AWS Region
C.The bucket policy blocks reading tags
D.The object was created without tags because of lifecycle rules
AnswerA

HeadObject returns object metadata and system headers, not the object's tag set, so the engineer sees no 'type=raw' tag despite it existing. GetObjectTagging is the API that retrieves tags, satisfying the expectation of confirming the object's tagging.

Why this answer

The head-object command does not return object tags; you must use the get-object-tagging command to retrieve tags. Option B is incorrect because the head-object command succeeds regardless of region, and region does not affect tag visibility. Option C is incorrect because bucket policies can deny access but do not prevent tags from being returned by head-object; they would affect get-object-tagging instead.

Option D is incorrect because lifecycle rules do not remove tags from objects; they may transition or expire objects but do not strip metadata.

593
MCQhard

A data engineer is designing a data ingestion pipeline for IoT sensor data. The sensors send JSON messages every second. The data must be available in Amazon S3 within 5 minutes and must be transformed (JSON to Parquet) before storage. Which combination of services meets these requirements?

A.Amazon Kinesis Data Streams with AWS Glue streaming ETL
B.Amazon Kinesis Data Firehose with data transformation and Parquet conversion
C.Amazon Kinesis Data Analytics with output to S3
D.Amazon S3 with S3 Event Notifications to AWS Lambda for transformation
AnswerB

Firehose buffers records and invokes a Lambda function for JSON-to-Parquet transformation before delivering to S3, meeting the five-minute latency requirement without managing clusters. Its native record format conversion writes Parquet directly, unlike Kinesis Data Streams, which only stores raw bytes.

Why this answer

Amazon Kinesis Data Firehose natively supports AWS Lambda-based record transformation and can convert the transformed JSON records to Parquet using its built-in format conversion before delivering to S3. Firehose buffers data and delivers within its configurable buffer interval (as low as 60 seconds), comfortably meeting the 5-minute SLA. This is the only option that combines ingestion, transformation, and Parquet conversion in a single managed service.

Exam trap

DEA-C01 often tests whether candidates know that Firehose has native Parquet conversion and Lambda transformation, versus assuming you must bolt on Glue or Lambda separately — leading them to pick the more complex option A.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Streams alone does not transform or convert to Parquet — you would need to add Glue streaming ETL, which adds complexity and typically has higher latency than Firehose's built-in transformation, and Data Streams requires custom consumers. Option C is wrong because Kinesis Data Analytics (now Managed Service for Apache Flink) performs stream processing but does not natively write Parquet to S3 without additional custom code, and it is overkill for simple JSON-to-Parquet conversion. Option D is wrong because S3 Event Notifications only fire after objects land in S3, meaning raw JSON would be stored first and transformation would be post-hoc, violating the 'transformed before storage' requirement and adding latency.

594
MCQhard

A company uses Amazon RDS for PostgreSQL with encryption at rest using AWS KMS. The company needs to share a database snapshot with a different AWS account. What must be done to allow the target account to restore the snapshot?

A.Copy the snapshot to the target account's region and share it
B.Create an IAM role in the source account that allows cross-account snapshot access
C.Share the snapshot and update the KMS key policy to allow the target account to use the key
D.Disable encryption on the snapshot before sharing
AnswerC

Sharing an encrypted RDS snapshot requires the target account to decrypt it, so the KMS key policy must grant that account kms:Decrypt and kms:CreateGrant. Without this key-policy change, the shared snapshot cannot be restored, because the target account lacks permission to use the customer-managed key.

Why this answer

Cross-account snapshot sharing of an encrypted snapshot requires both sharing the snapshot and granting the target account permission to use the KMS key via the key policy. Option A is incorrect because copying does not grant the necessary key access. Option B is incorrect because IAM roles are not used for this purpose; KMS key policies are the mechanism for cross-account access.

Option D is incorrect because encryption cannot be disabled on an existing encrypted snapshot.

595
MCQmedium

A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another. The security team requires that all data be encrypted at rest and that the job use a customer-managed KMS key for encryption. The engineer configures the job with `--encryption-type sse-kms` and a KMS key ID. However, the job fails with an access denied error when writing to S3. The IAM role used by the Glue job has permissions to read and write the S3 buckets but has no KMS permissions. Which additional IAM permissions are required for the Glue job role to successfully write encrypted data?

A.kms:Decrypt and kms:DescribeKey on the KMS key
B.kms:ReEncryptFrom and kms:ReEncryptTo on the KMS key
C.kms:CreateGrant and kms:ListKeys on the KMS key
D.kms:Encrypt and kms:GenerateDataKey on the KMS key
AnswerD

When Glue writes data to S3 using SSE-KMS, the service must call AWS KMS to generate a data key and encrypt it. The Glue job role needs kms:Encrypt and kms:GenerateDataKey permissions on the customer-managed KMS key. Without these, the write operation fails with access denied. These permissions allow the job to encrypt the data key and use it for SSE-KMS.

Why this answer

To write data to Amazon S3 with SSE-KMS, the Glue job role must have permissions to use the KMS key for encryption. Specifically, kms:Encrypt and kms:GenerateDataKey are required to create and encrypt the data key used for server-side encryption. Without these, the S3 PutObject operation fails with access denied.

The other options either relate to reading (Decrypt) or administrative tasks (CreateGrant, ReEncrypt) that do not satisfy the encryption requirement for writing.

Exam trap

The trap here is assuming that S3 write permissions alone are sufficient for SSE-KMS encryption, forgetting that the KMS key policy and IAM role must also grant encryption permissions.

596
MCQhard

A data engineer maintains an Amazon DynamoDB table that stores device telemetry. The table uses a partition key of deviceId and a sort key of timestamp, with on-demand capacity mode. A new fleet of devices writes data with a deviceId pattern that hashes to a small number of partitions, and the engineer observes throttling on writes even though consumed capacity is well below any configured limit. Which change addresses the root cause?

A.Create a global secondary index on timestamp and direct all writes through that index to distribute the load.
B.Switch the table from on-demand to provisioned capacity mode and raise the write capacity units substantially.
C.Redesign the key schema to use a higher-cardinality partition key, such as a composite of deviceId and a shard suffix, while keeping timestamp as the sort key.
D.Enable DynamoDB Streams on the table so that write operations are buffered and replayed during peak periods.
AnswerC

DynamoDB partitions data by the hash of the partition key, so a low-cardinality key concentrates writes onto a few partitions and triggers throttling regardless of overall table capacity. Introducing a shard suffix spreads writes across many partitions while the sort key still allows efficient range queries per device, directly removing the hot-partition bottleneck.

Why this answer

Throttling at low consumed capacity on a DynamoDB table almost always indicates a hot partition caused by an uneven partition key. Because data is distributed by the hash of the partition key, a key with few distinct values funnels traffic into a small number of partitions. Adding a shard suffix to the partition key spreads the writes while preserving sort-key range queries.

Exam trap

The trap here is treating throttling as a capacity shortage and raising throughput, when the real cause is a partition key whose hash space is too small.

597
MCQhard

A data engineer is designing a data lake on Amazon S3 and needs to catalog data using the AWS Glue Data Catalog. The data is stored in Parquet format, partitioned by year/month/day. The engineer wants to query the data using Amazon Athena and ensure that partition pruning occurs to minimize query costs. Which action should the engineer take?

A.Use AWS Glue crawlers to automatically discover the schema and partitions, then run MSCK REPAIR TABLE in Athena.
B.Use AWS Glue ETL jobs to write data to S3 in a non-partitioned layout and rely on Athena's predicate pushdown.
C.Manually define the table in AWS Glue Data Catalog with partition keys and use Athena's ALTER TABLE ADD PARTITION for each partition.
D.Configure the table in AWS Glue Data Catalog with partition projection using the appropriate storage descriptor and table properties.
AnswerD

Partition projection allows Athena to calculate partition locations dynamically based on table properties, eliminating the need to manually add partitions or run crawlers. This enables efficient partition pruning and reduces query costs, especially for highly partitioned datasets like year/month/day.

Why this answer

Partition projection in AWS Glue Data Catalog enables Athena to dynamically compute partition locations, avoiding the need to manage partitions manually. It supports efficient partition pruning, reducing the amount of data scanned and lowering query costs. This is the recommended approach for highly partitioned datasets with a predictable structure.

Exam trap

The trap here is assuming that running MSCK REPAIR TABLE or using Glue crawlers is sufficient for partition pruning, when in fact partition projection provides a more scalable and cost-effective solution.

598
MCQhard

A data engineer is building a streaming pipeline using Amazon Kinesis Data Streams and AWS Lambda. The Lambda function processes records and writes to Amazon DynamoDB. The engineer notices that the Lambda function is throttled during high traffic. Which action should the engineer take to reduce throttling?

A.Increase the Lambda function timeout
B.Disable retries on the Lambda function
C.Increase the number of shards in the Kinesis data stream
D.Use an Amazon SQS queue as an intermediate buffer
AnswerC

More shards allow more Lambda concurrent executions, reducing throttling.

Why this answer

Increasing the number of shards in the Kinesis data stream increases the overall throughput of the stream, which allows the Lambda event source mapping to poll more shards concurrently. Each shard is processed by one Lambda invocation at a time, so more shards mean more concurrent Lambda executions, reducing the per-invocation load and the likelihood of throttling.

Exam trap

The DEA-C01 exam often tests the misconception that throttling is caused by Lambda function performance (timeout or retries) rather than the stream's shard count, leading candidates to choose options that affect execution duration or error handling instead of scaling the source.

How to eliminate wrong answers

Option A is wrong because increasing the Lambda function timeout does not reduce throttling; it only allows the function to run longer, which does not affect the rate at which Lambda invokes the function. Option B is wrong because disabling retries on the Lambda function would cause records to be dropped or sent to a dead-letter queue, but it does not prevent throttling; throttling occurs when the concurrent execution limit is reached, not from retries. Option D is wrong because using an Amazon SQS queue as an intermediate buffer would decouple the stream from Lambda but does not directly address the root cause of throttling, which is insufficient shard count to handle the incoming data volume; SQS would add latency and complexity without increasing concurrency.

599
MCQhard

A data engineer is troubleshooting an AWS Glue ETL job that suddenly started failing with 'An error occurred while calling o103.pyWriteDynamicFrame. Unknown error'. The job writes data to an Amazon Redshift table. Which step should the engineer take FIRST?

A.Recreate the Redshift table with a different distribution style.
B.Test the job with a small sample dataset to isolate the issue.
C.Update the Redshift JDBC driver version in the Glue job.
D.Review the job's CloudWatch Logs for detailed error messages.
AnswerD

The error message 'An error occurred while calling o103.pyWriteDynamicFrame. Unknown error' is generic and does not specify the root cause. The first troubleshooting step should be to review the job's CloudWatch Logs, which provide detailed error messages, stack traces, and other diagnostic information.

Why this answer

The error message 'An error occurred while calling o103.pyWriteDynamicFrame. Unknown error' is a generic wrapper thrown by the Glue DynamicFrame writer when the underlying cause is not surfaced. The first and most reliable step is to inspect the job's CloudWatch Logs, where the full stack trace, Redshift error code, and driver-level messages are recorded.

This aligns with standard AWS troubleshooting guidance: always check logs before making changes. Reviewing logs is non-destructive, fast, and directly reveals the root cause (e.g., permission issue, connection timeout, schema mismatch).

Exam trap

DEA-C01 often tests the principle of 'logs first' in troubleshooting scenarios, where candidates are tempted to jump to speculative fixes like changing configurations or drivers instead of gathering diagnostic data.

How to eliminate wrong answers

Option A is wrong because changing the distribution style is a speculative fix that does not address the immediate need to diagnose the error; it could also introduce performance regressions. Option B is wrong because testing with a small dataset might reproduce the error but does not provide the detailed diagnostic information that logs offer, and it consumes additional time and resources. Option C is wrong because updating the JDBC driver is a potential solution only after confirming a driver incompatibility, which the logs would reveal; doing it first is premature and may not resolve the issue.

600
MCQeasy

A company needs to ingest streaming data from multiple sources and store it in Amazon S3. The data volume is up to 5 GB per hour. What is the MOST cost-effective ingestion service?

A.AWS Glue
B.Amazon Kinesis Data Streams
C.Amazon Kinesis Data Firehose
D.AWS Lambda
AnswerC

Kinesis Data Firehose buffers and delivers streaming records directly into Amazon S3 without managing consumers, and its pricing suits 5 GB per hour. This satisfies the stem's cost-effectiveness constraint while meeting the ingestion and S3 delivery requirement.

Why this answer

Amazon Kinesis Data Firehose is the most cost-effective service for ingesting streaming data into Amazon S3 at 5 GB/hour. It is fully managed, automatically scales, and charges only for data ingested (per GB), with no upfront provisioning. While Amazon Kinesis Data Streams (KDS) may have lower throughput cost for steady loads, it requires manual shard management and typically needs additional components (e.g., Lambda functions) to deliver data to S3, increasing operational overhead and total cost.

AWS Glue is a batch ETL service, not designed for streaming. AWS Lambda is a compute service and would require custom code and scaling logic, making it more expensive and complex for this use case. Therefore, Kinesis Data Firehose provides the simplest and most cost-effective solution for streaming data ingestion directly to S3.

Page 7

Page 8 of 18

Page 9