Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 976–1050

1321 questions total · 18pages · All types, answers revealed

Page 13

Page 14 of 18

Page 15
976
MCQhard

A data engineer is troubleshooting an ETL job that reads from an S3 bucket encrypted with SSE-KMS. The job is failing with an error indicating that the IAM role does not have permission to decrypt the data. What is the most likely missing permission?

A.kms:GenerateDataKey
B.s3:ListBucket
C.kms:Decrypt
D.s3:GetObject
AnswerC

SSE-KMS requires the caller to hold kms:Decrypt on the customer-managed key before S3 can return the object. The role's S3 permissions are irrelevant here; the missing KMS grant is what blocks the ETL job's reads.

Why this answer

When an S3 object is encrypted with SSE-KMS, reading the object requires two sets of permissions: s3:GetObject on the object and kms:Decrypt on the KMS key used to encrypt it. The error explicitly states the IAM role lacks permission to decrypt, so the missing permission is kms:Decrypt. Without it, S3 cannot call KMS to decrypt the data key, and the GetObject call fails with AccessDenied.

Exam trap

The trap is assuming that s3:GetObject alone is sufficient to read SSE-KMS encrypted objects, ignoring the separate KMS permission required for decryption.

How to eliminate wrong answers

Option A is wrong because kms:GenerateDataKey is needed for writing (PutObject) with SSE-KMS, not for reading existing encrypted objects. Option B is wrong because s3:ListBucket controls the ability to list objects in a bucket, which is unrelated to decryption and would produce a different error. Option D is wrong because s3:GetObject is the permission to read the object itself; if it were missing, the error would be about S3 access, not KMS decryption.

977
MCQeasy

A company wants to audit all changes to IAM policies in their AWS account. Which AWS service should be used to record these changes for compliance purposes?

A.Amazon CloudWatch Logs
B.AWS Config
C.AWS CloudTrail
D.Amazon S3
AnswerC

AWS CloudTrail records API activity, capturing every IAM policy change as a management event with the caller's identity, timestamp and source IP. This satisfies the audit requirement by providing an immutable, queryable log of who modified which policy and when, which CloudWatch metrics or Config rules alone cannot deliver.

Why this answer

AWS CloudTrail records API calls, including IAM policy changes. AWS Config records resource configurations but not all API calls. CloudWatch Logs can store logs but does not record API calls itself.

S3 is the destination for logs, not the recording service.

978
MCQmedium

A data engineer is using AWS Lake Formation to manage fine-grained access to a data lake in Amazon S3. The engineer grants a data analyst SELECT permission on a table but wants to ensure that the analyst cannot access columns containing sensitive data such as social security numbers. The table is registered in the AWS Glue Data Catalog. Which Lake Formation feature should the engineer use to restrict access to specific columns?

A.Row-level security using filter expressions in Lake Formation.
B.S3 bucket policies with prefix-based restrictions.
C.AWS Glue Data Catalog resource policies.
D.Column-level security in Lake Formation permissions.
AnswerD

Lake Formation supports column-level security by allowing you to grant or deny access to specific columns within a table. When granting SELECT on a table, you can exclude sensitive columns, ensuring the analyst can query only approved columns. This directly meets the requirement without duplicating data.

Why this answer

Lake Formation column-level security allows granting SELECT on a subset of columns, effectively hiding sensitive columns from unauthorized users. Row-level security filters rows, not columns. Glue Data Catalog policies and S3 bucket policies do not provide column-level granularity for query engines.

Exam trap

The trap here is mixing up row-level and column-level security features in Lake Formation, or assuming that S3 or Glue policies can enforce column-level access.

979
MCQeasy

A data engineer needs to audit all changes to IAM policies in an AWS account. Which AWS service should be used?

A.AWS CloudTrail
B.AWS Config
C.AWS Organizations
D.Amazon CloudWatch Logs
AnswerA

CloudTrail records all API activity for auditing.

Why this answer

WS CloudTrail because it records API calls made in the account, including changes to IAM policies. AWS Config tracks resource configuration changes, not API calls. AWS Organizations is for multi-account management.

Amazon CloudWatch Logs is for log storage and monitoring, not auditing API calls.

980
MCQhard

A company uses AWS Glue to run ETL jobs that process data from Amazon S3 and load into Amazon Redshift. The jobs have recently started failing with 'Out of Memory' errors. The data volume has increased 3x in the past month. Which is the MOST effective solution to resolve this issue without redesigning the job?

A.Use Amazon Athena instead of Glue for the transformation.
B.Increase the number of Glue workers (DPUs) for the job.
C.Rewrite the job to use Spark SQL instead of PySpark.
D.Increase the number of partitions in the input S3 data.
AnswerB

Increasing the number of Glue workers (DPUs) adds parallel executors and memory capacity, letting the existing job process the tripled data volume without code changes. Vertical scaling of worker type alone is less effective than adding workers for memory-bound shuffles.

Why this answer

Glue ETL jobs run on a Spark cluster sized by the number of DPUs (workers). Out-of-memory errors during a 3x data volume increase indicate the existing worker count cannot hold the partitions and shuffle data in memory. Increasing the number of Glue workers adds executor memory and parallelism, which is the most direct fix without redesigning the job.

Exam trap

DEA-C01 often tests the misconception that changing the query language or input partitioning fixes OOM, when the direct lever for Spark memory is the number and size of Glue workers.

How to eliminate wrong answers

Option A is wrong because switching to Athena is a redesign of the processing architecture and is not a drop-in fix for an existing Glue job. Option C is wrong because rewriting PySpark as Spark SQL does not inherently reduce memory pressure and still requires code changes. Option D is wrong because increasing S3 partitions can help parallelism but does not guarantee more memory per executor and may even increase overhead if not matched with more workers.

981
MCQmedium

A data engineer needs to load data from an on-premises Oracle database to Amazon S3 daily. The table is 500 GB and grows by 50 MB per day. The load must capture only new and changed rows since the last run. Which solution is MOST cost-effective and requires the least maintenance?

A.Write a custom Python script on EC2 to query the Oracle redo logs and upload to S3
B.Export the entire table to CSV daily using a script and upload to S3
C.Use AWS Glue ETL job with a JDBC connection and a timestamp filter
D.Use AWS Database Migration Service (DMS) with ongoing replication (CDC)
AnswerD

DMS supports CDC and can capture only changes, minimizing cost and effort.

Why this answer

AWS DMS with ongoing replication (CDC) is the most cost-effective and low-maintenance solution because it continuously captures only new and changed rows from the Oracle source using its built-in CDC mechanism (reading redo logs), without requiring custom scripting or full table exports. It automatically handles schema changes, resumability, and incremental loading to S3, minimizing operational overhead and data transfer costs.

Exam trap

The trap here is that candidates often choose AWS Glue with a timestamp filter (Option C) because it seems simpler, but they overlook that Glue still performs a full table scan via JDBC to apply the filter, which is inefficient for large tables and does not provide true CDC from redo logs, unlike DMS's native log-based replication.

How to eliminate wrong answers

Option A is wrong because writing a custom Python script on EC2 to parse Oracle redo logs is complex to implement, requires deep Oracle internals knowledge, and demands ongoing maintenance for log format changes and error handling, making it neither cost-effective nor low-maintenance. Option B is wrong because exporting the entire 500 GB table daily to CSV is extremely inefficient, wastes significant compute and network resources, and incurs high S3 storage costs for unchanged data, failing the 'only new and changed rows' requirement. Option C is wrong because while AWS Glue with a timestamp filter can capture incremental changes, it requires the source table to have a reliable, monotonically increasing timestamp column and still performs a full JDBC scan of the table to filter rows, which is inefficient for a 500 GB table and does not natively support change data capture from redo logs.

982
MCQhard

Refer to the exhibit. A CloudFormation stack outputs the Glue job name and S3 bucket names. The Glue job transforms CSV files from the raw bucket to Parquet in the processed bucket. However, the Glue job is failing with an error that it cannot write to the processed bucket. What is the most likely cause?

A.The Glue job does not have permission to write to the processed bucket
B.The raw data bucket is in a different region
C.The Glue job is not using the correct worker type
D.The Glue job is using an incorrect file format
AnswerA

CloudFormation outputs merely expose values; they grant no IAM permissions. The Glue job's execution role therefore lacks s3:PutObject on the processed bucket, producing the access-denied write failure. The fix is attaching a policy permitting writes to that bucket.

Why this answer

The most likely cause is that the Glue job's IAM role lacks the necessary permissions (e.g., s3:PutObject, s3:ListBucket) on the processed bucket. AWS Glue jobs require an IAM role with policies that grant write access to the target S3 bucket; without these permissions, the job fails with a write error. This is a common misconfiguration when the role is scoped only to read from the raw bucket.

Exam trap

The DEA-C01 exam often tests the misconception that S3 write failures are caused by region mismatches or file format issues, but the actual trap is that candidates overlook the IAM permission layer and attribute the error to non-permission factors like worker type or data format.

How to eliminate wrong answers

Option B is wrong because cross-region access to S3 buckets is fully supported and does not cause a write permission error; the error message specifically indicates a write failure, not a connectivity or region mismatch. Option C is wrong because the worker type (e.g., G.1X, G.2X) affects memory and compute capacity, not S3 write permissions; an incorrect worker type would cause performance or OOM issues, not a bucket write error. Option D is wrong because using an incorrect file format (e.g., specifying Parquet when the output is CSV) would cause a format conversion error, not a permission-denied write error; the error message explicitly states it cannot write to the bucket, pointing to an access control issue.

983
Multi-Selecteasy

A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data must be transformed in real-time using a custom Lambda function. Which TWO steps are required to enable this? (Choose TWO)

Select 2 answers
A.Configure Kinesis Data Firehose to use a Lambda function for data transformation
B.Ensure the Lambda function returns the transformed records in the correct format
C.Create a Kinesis Data Analytics application to transform the data
D.Write the transformation logic directly in the Firehose delivery stream configuration
E.Use Kinesis Data Streams as the source for Firehose
AnswersA, B

Firehose only invokes a Lambda function for transformation when the delivery stream is explicitly configured with that function's Amazon Resource Name under its processing parameters. Without this configuration, records pass through untransformed, so the real-time transformation requirement cannot be met.

Why this answer

Option A is correct because enabling real-time transformation in Kinesis Data Firehose requires configuring the delivery stream to invoke a Lambda function as its processing step, which Firehose then calls synchronously for each batch of records. Option B is correct because the Lambda function must return the transformed records in the exact format Firehose expects — a JSON object containing the processed records with fields such as recordId, result (Ok, Dropped, or ProcessingFailed), and data (base64-encoded) — otherwise Firehose cannot continue delivery. Option C is wrong because Kinesis Data Analytics is a separate service for SQL/Flink stream analytics, not the mechanism Firehose uses for Lambda-based transformation.

Option D is wrong because Firehose does not execute inline transformation logic; it only supports invoking a Lambda function for that purpose. Option E is wrong because using Kinesis Data Streams as a source is unrelated to enabling Lambda transformation and is not a required step.

Exam trap

DEA-C01 often tests the Firehose transformation trap: candidates assume transformation logic is written inside the Firehose configuration or that Kinesis Data Analytics is required, when in fact Firehose delegates to a Lambda function that must return a specific JSON envelope.

984
MCQmedium

A healthcare company stores patient records in an Amazon S3 bucket and uses AWS Lake Formation to manage access for multiple analytics teams. The compliance team requires that any column containing patient identifiers be masked by default for all users except a privileged data steward role. Which Lake Formation feature should the data engineer implement to meet this requirement?

A.Enable S3 Object Lock in governance mode on the bucket to prevent unauthorized access to sensitive objects.
B.Use AWS Glue DataBrew to create a masking recipe that anonymizes the sensitive columns before granting access.
C.Configure a data filter that excludes sensitive columns from the table definition.
D.Create a Lake Formation tag-based access control policy that attaches LF-Tags to the sensitive columns and grants access only to the steward role.
AnswerD

Lake Formation tag-based access control (LF-TBAC) allows you to attach LF-Tags to databases, tables, and columns, and then grant permissions based on those tags. By tagging sensitive columns and granting access only to the steward role, all other users are denied access to those columns by default. This satisfies the requirement to mask by default while allowing the steward full access.

Why this answer

Lake Formation tag-based access control enables attribute-based permissions where LF-Tags on columns determine access. Tagging sensitive columns and granting only the steward role access ensures all other users are denied those columns by default, effectively masking them. This is a native Lake Formation governance feature that integrates with analytics services and provides fine-grained control without duplicating data.

Exam trap

The trap here is assuming that data filters in Lake Formation can mask column values, when they actually restrict rows or hide columns entirely.

985
Multi-Selectmedium

A company uses Amazon Redshift for analytics. The data engineering team wants to improve query performance for frequently used aggregate queries. Which TWO actions would help achieve this?

Select 2 answers
A.Increase the number of WLM query queues
B.Use distribution keys to collocate data on the same node slices
C.Run the VACUUM command to reclaim space from deleted rows
D.Define appropriate sort keys on the tables
E.Increase the number of nodes in the cluster
AnswersB, D

Distribution keys determine which node slice stores each row, so collocating joined or aggregated rows on the same slice lets Redshift perform local joins and partial aggregation, cutting data movement across the network during aggregate queries.

Why this answer

Option B is correct because choosing an appropriate distribution key collocates matching rows on the same node slices, so joins and aggregations can be processed locally without expensive data redistribution (broadcast or shuffle) across the cluster, directly speeding up frequently used aggregate queries. Option D is correct because defining appropriate sort keys physically orders data on disk by the key columns, enabling zone maps to skip irrelevant blocks and allowing efficient range-restricted scans and merge joins, which reduces the data read for aggregate queries. Option A is not correct because adding WLM query queues only changes concurrency and memory allocation among query groups; it does not by itself make an individual aggregate query faster.

Option C is not correct because VACUUM reclaims space from deleted rows and re-sorts data, which is a maintenance operation rather than a design change that improves aggregate query performance. Option E is not correct because adding nodes increases cluster capacity and parallelism but does not address the underlying data layout, so poorly distributed or unsorted tables can still cause slow aggregate queries.

Exam trap

The trap here is that candidates often confuse VACUUM (which reclaims space) with performance optimization for queries, or assume adding nodes always improves query speed without considering the overhead of data redistribution.

986
MCQmedium

A data engineer is using Amazon Redshift and needs to improve query performance for a large fact table that is frequently joined with a much smaller dimension table. The engineer wants to minimize data movement during joins. Which distribution style should be used for the dimension table?

A.DISTSTYLE AUTO
B.DISTSTYLE KEY
C.DISTSTYLE ALL
D.DISTSTYLE EVEN
AnswerC

DISTSTYLE ALL replicates the entire dimension table to every compute node. When joining with a large fact table, each node already has a copy of the dimension table, eliminating the need to redistribute the dimension data during the join. This minimizes data movement and improves join performance, which is ideal for small dimension tables.

Why this answer

For a small dimension table that is frequently joined with a large fact table, DISTSTYLE ALL replicates the dimension table to all nodes. This ensures that each node has a local copy, eliminating the need to broadcast or redistribute the dimension data during joins. As a result, data movement is minimized, and join performance is significantly improved.

Exam trap

The trap here is assuming that automatic distribution or distributing on the join key is always best, when for small dimension tables, replicating the entire table with DISTSTYLE ALL is often more efficient.

987
MCQmedium

A data engineer is building an AWS Glue ETL job that reads records from an Amazon Kinesis Data Stream and writes them to Amazon S3 in Parquet format. The job must checkpoint its progress so that it can resume without reprocessing data after a failure. Which AWS Glue mechanism should the engineer configure to track the stream position?

A.Use AWS Glue Studio's auto-generated script and enable the 'retry on failure' option in the job's trigger.
B.Enable job bookmarks with the transformation_ctx parameter set on the create_data_frame_from_options call.
C.Set the Glue job's MaxConcurrentRuns parameter to 1 and let the job restart from the beginning on failure.
D.Configure a Kinesis Data Stream consumer with enhanced fan-out and rely on the stream's 24-hour retention to replay records.
AnswerB

Job bookmarks persist the state of the last processed record per source. When reading from Kinesis, Glue stores the stream sequence number and shard information, so a re-run resumes from the checkpoint. Setting transformation_ctx uniquely identifies the source in the bookmark store, preventing conflicting state when a job reads multiple streams.

Why this answer

Glue job bookmarks maintain state between runs, storing the last processed Kinesis sequence number per shard. With transformation_ctx set on the read operation, the job can resume from where it stopped after a failure, avoiding duplicate writes to S3. The other options address throughput, concurrency, or retry orchestration, none of which track record-level progress.

Exam trap

The trap here is assuming that Kinesis stream retention or retry triggers provide checkpointing, when only Glue job bookmarks persist the last processed position.

988
MCQmedium

A company uses Amazon S3 to store large CSV files and runs Amazon Athena queries on them. The queries are becoming slower as data grows. A data engineer suggests converting the files to Apache Parquet format and partitioning the data. What is the primary benefit of converting to Parquet?

A.Parquet allows schema evolution without rewriting files.
B.Parquet supports nested data structures that CSV cannot.
C.Parquet stores data in a columnar format, reducing the amount of data scanned per query.
D.Parquet is compressed by default, reducing storage costs.
AnswerC

Parquet's columnar layout lets Athena read only the columns referenced in each query rather than every field in each row, so far less data is scanned from S3. Partitioning prunes whole prefixes, but the format change itself is what satisfies the stem's primary benefit of reduced scan volume.

Why this answer

Parquet is a columnar storage format that stores data by columns rather than rows. When Athena queries only a subset of columns, it can read just those columns from disk, drastically reducing the amount of data scanned per query. This directly addresses the performance slowdown because Athena charges by data scanned, and less scanning means faster queries and lower costs.

Exam trap

The trap here is that candidates confuse the general benefits of Parquet (compression, schema evolution, nested data) with the primary performance benefit for Athena, which is columnar pruning reducing scanned data.

How to eliminate wrong answers

Option A is wrong because Parquet does support schema evolution (e.g., adding columns) but this is not its primary benefit for query performance; schema evolution is a feature of many formats and not unique to Parquet's columnar nature. Option B is wrong because while Parquet does support nested data structures (like structs and arrays), CSV does not, but this is a data modeling advantage, not the primary performance benefit for large-scale analytics queries. Option D is wrong because Parquet is not compressed by default; compression is configurable (e.g., Snappy, Gzip, Zstd) and while it reduces storage costs, the primary benefit for query speed is columnar pruning, not compression.

989
Multi-Selectmedium

A company is building a data lake on AWS and must encrypt data at rest. Which services can provide server-side encryption for data stored in Amazon S3? (Choose TWO.)

Select 2 answers
A.SSE-S3
B.SSL/TLS
C.AWS SDK client-side encryption
D.AWS CloudHSM
E.SSE-KMS
AnswersA, E

SSE-S3 applies AES-256 encryption with keys fully managed and rotated by AWS, requiring no key configuration from the engineer. It satisfies the data-at-rest encryption requirement for S3 objects while removing customer key management overhead entirely.

Why this answer

SSE-S3 (Option A) is correct because Amazon S3 server-side encryption with S3-managed keys (AES-256) encrypts objects at rest automatically, with AWS managing the key material and rotation. SSE-KMS (Option E) is also correct because it performs server-side encryption at rest using AWS KMS customer master keys (CMKs), giving you control over key policies, auditing, and rotation. Both are S3 server-side encryption modes applied after data reaches S3, satisfying the data-at-rest requirement.

SSL/TLS (Option B) is wrong because it only encrypts data in transit between the client and S3, not at rest. AWS SDK client-side encryption (Option C) is wrong because it encrypts data before it is sent to S3, so it is client-side, not server-side. AWS CloudHSM (Option D) is wrong because it is a dedicated hardware security module service for key storage and cryptographic operations, not an S3 server-side encryption option by itself.

Exam trap

DEA-C01 often tests whether candidates confuse encryption at rest (SSE-S3, SSE-KMS) with encryption in transit (SSL/TLS) or client-side encryption, and may mistakenly select CloudHSM as a direct S3 encryption option.

990
MCQhard

A company ingests JSON data from an S3 bucket into a Glue ETL job. The data contains nested structures and arrays. The team wants to flatten the data into a tabular format for analysis in Athena. Which Glue transformation is appropriate?

A.Map
B.Relationalize
C.Filter
D.DropNullFields
AnswerB

Relationalize converts nested JSON and arrays into a set of flat, related tables, which is precisely what Athena needs for tabular SQL queries. Other transforms such as ApplyMapping only rename or retype fields and cannot unnest arrays, so they would leave the nested structures intact.

Why this answer

The Relationalize transformation is specifically designed to flatten nested JSON and arrays into a tabular format suitable for Athena. Option A (Map) applies a function to each record but does not flatten structures. Option C (Filter) selects rows based on a condition.

Option D (DropNullFields) removes null fields but does not address nested structures.

991
MCQmedium

A data engineer must give an Amazon Redshift cluster the ability to load data from an Amazon S3 bucket using the COPY command. The security team prohibits embedding long-term AWS credentials in SQL and requires that access be revoked automatically when the cluster is deleted. The S3 bucket is encrypted with SSE-KMS using a customer managed key. Which approach should the data engineer use?

A.Configure the cluster with a database user that has a password stored in AWS Systems Manager Parameter Store and grant that user access to the S3 bucket through a bucket ACL.
B.Use the COPY command with the ACCESS_KEY_ID and SECRET_ACCESS_KEY parameters populated from an IAM role's temporary credentials obtained by calling AssumeRole from an external application.
C.Create an IAM user with an access key, store the key in AWS Secrets Manager, and pass the secret ARN to the COPY command using the CREDENTIALS clause.
D.Attach an IAM role to the Redshift cluster, grant the role s3:GetObject on the bucket and kms:Decrypt on the customer managed key, and reference the role ARN in the COPY command.
AnswerD

Attaching an IAM role to the cluster provides temporary credentials that Redshift assumes automatically, so no long-term keys appear in SQL. Granting s3:GetObject and kms:Decrypt allows the COPY to read SSE-KMS encrypted objects. Because the role is attached to the cluster, deleting the cluster removes the role association, satisfying automatic revocation.

Why this answer

Attaching an IAM role to the Amazon Redshift cluster lets the COPY command assume temporary credentials without any long-term keys in SQL. The role needs s3:GetObject on the bucket and kms:Decrypt on the customer managed key to read SSE-KMS encrypted objects. Because the role is associated with the cluster, deleting the cluster removes that association, so access is revoked automatically as the security team requires.

Exam trap

The trap here is assuming the COPY command can accept arbitrary credential parameters, when the supported credential-free path is an IAM role attached to the cluster referenced by ARN.

992
MCQhard

A data engineer notices that an Amazon Redshift cluster’s storage usage is increasing rapidly due to many UPDATE and DELETE operations. The engineer needs to reclaim storage space and improve query performance. Which action should be taken?

A.Run VACUUM command
B.UNLOAD the table to S3 and reload
C.Increase cluster node count
D.Run ANALYZE command
AnswerA

VACUUM reclaims space from deleted rows and re-sorts unsorted regions, restoring sequential scan performance degraded by frequent UPDATE and DELETE operations. This directly addresses the stem's storage growth and query-performance constraints on the Redshift cluster.

Why this answer

The VACUUM command in Amazon Redshift reclaims disk space occupied by deleted or updated rows and re-sorts the data according to the table's sort keys. This directly addresses the storage increase from UPDATE/DELETE operations and improves query performance by restoring the physical order of rows, which reduces the number of blocks scanned.

Exam trap

The trap here is that candidates confuse ANALYZE with VACUUM, thinking updating statistics will also reclaim storage, when in fact ANALYZE only refreshes metadata for the query optimizer and has no effect on physical storage.

How to eliminate wrong answers

Option B is wrong because unloading the table to S3 and reloading is a heavy, manual process that does not reclaim space in place and can be avoided with a simple VACUUM; it also incurs additional S3 costs and time. Option C is wrong because increasing the cluster node count adds more storage and compute capacity but does not reclaim the existing wasted space from deleted rows, and it may not improve performance if the underlying data is fragmented. Option D is wrong because the ANALYZE command only updates table statistics for the query planner, it does not reclaim storage space or physically reorganize data affected by UPDATE/DELETE operations.

993
MCQeasy

A data engineer needs to store semi-structured data (JSON logs) from thousands of IoT devices. The data must be schema-less, highly scalable, and support low-latency queries by device ID and timestamp. Which AWS service should the engineer use?

A.Amazon RDS for PostgreSQL
B.Amazon Redshift
C.Amazon DynamoDB
D.Amazon S3
AnswerC

DynamoDB stores JSON as native map and list attributes without a fixed schema, scales horizontally, and a composite partition key of device ID plus sort key on timestamp delivers low-latency item queries — matching the schema-less, scalable, low-latency constraints exactly.

Why this answer

Amazon DynamoDB is the correct choice because it is a fully managed NoSQL key-value and document database that natively supports semi-structured JSON data, schema-less design, and automatic scaling. Its partition key (device ID) and sort key (timestamp) enable low-latency, single-millisecond queries by device ID and timestamp, making it ideal for high-throughput IoT log ingestion.

Exam trap

The trap here is that candidates often confuse Amazon S3's ability to store JSON files with the ability to query them efficiently, overlooking that S3 lacks native indexing and low-latency query support, which DynamoDB provides through its key-value access pattern.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for PostgreSQL is a relational database with a fixed schema, requiring predefined tables and indexes for JSON data, which cannot handle schema-less IoT logs at scale without manual sharding or performance tuning. Option B is wrong because Amazon Redshift is a columnar data warehouse optimized for analytical queries on structured data, not for low-latency point queries by device ID and timestamp, and its schema-on-write model conflicts with schema-less requirements. Option D is wrong because Amazon S3 is an object store that can store JSON logs but lacks native indexing and low-latency query capabilities; querying by device ID and timestamp would require scanning or external services like Athena, adding latency and complexity.

994
MCQeasy

A company stores application logs in an Amazon S3 bucket. A compliance policy states that log objects must be retained for exactly 90 days and then permanently deleted, and that no one, including administrators, should be able to delete them earlier. The data engineer must enforce this with the least effort. What should the engineer do?

A.Create an S3 Lifecycle rule that transitions objects to S3 Glacier Deep Archive after 90 days.
B.Apply an S3 Object Lock retention period of 90 days in compliance mode to the bucket, and configure a lifecycle rule to expire objects after 90 days.
C.Use AWS Backup to create a vault with a 90-day retention and a vault lock in compliance mode.
D.Enable S3 Versioning and add a bucket policy that denies s3:DeleteObject to all principals.
AnswerB

S3 Object Lock in compliance mode prevents any user, including the root user, from deleting or overwriting an object version until the retention period expires. Setting a 90-day retention plus a lifecycle expiration rule enforces both the immutability and the automatic deletion after 90 days, satisfying the policy with minimal ongoing effort.

Why this answer

S3 Object Lock in compliance mode provides WORM protection that even the root user cannot override before the retention period ends, which satisfies the immutability requirement. Pairing it with a lifecycle rule that expires objects after 90 days ensures automatic permanent deletion at the required time. Together they enforce the compliance policy with minimal operational effort.

Exam trap

The trap here is treating versioning or lifecycle transitions as immutability controls, when only Object Lock in compliance mode prevents deletion by any principal during the retention period.

995
MCQmedium

A company runs a Redshift cluster and notices that query performance has degraded over time. The data engineer suspects that table statistics are stale. What should the engineer do to improve query performance?

A.Rebuild the tables by using CREATE TABLE AS
B.Increase the number of slices in the cluster
C.Run the ANALYZE command on the tables
D.Run the VACUUM command on the tables
AnswerC

Stale statistics cause the Redshift query optimiser to misjudge cardinality, producing poor join orders and distribution choices. Running ANALYZE refreshes table statistics so the optimiser generates a more efficient plan, directly addressing the degraded performance described in the stem.

Why this answer

Stale table statistics cause the Redshift query optimizer to generate suboptimal execution plans, leading to degraded query performance. Running the ANALYZE command updates these statistics, allowing the optimizer to make better decisions about join order, distribution, and data scan strategies. This directly addresses the root cause of performance degradation over time.

Exam trap

The trap here is confusing the VACUUM command (which reorganizes physical storage) with the ANALYZE command (which updates query optimizer metadata), leading candidates to choose VACUUM when stale statistics are the actual culprit.

How to eliminate wrong answers

Option A is wrong because rebuilding tables with CREATE TABLE AS (CTAS) does not update statistics; it creates a new table that still requires an explicit ANALYZE to populate its statistics, and it is an unnecessarily heavy operation for fixing stale stats. Option B is wrong because increasing the number of slices in the cluster requires resizing the cluster (e.g., adding nodes or changing node types), which is a disruptive, costly operation that does not address stale statistics; query performance degradation from stale stats is not resolved by adding more slices. Option D is wrong because the VACUUM command reclaims disk space and sorts rows to maintain physical data organization, but it does not update table statistics; stale statistics persist after VACUUM, so the optimizer remains uninformed.

996
MCQhard

A company runs a daily batch ETL job using AWS Glue. The job processes 500 GB of data from Amazon RDS to Amazon S3. The job currently uses a single DPU and takes 6 hours to complete. The team wants to reduce runtime to under 1 hour without increasing costs significantly. Which approach should they use?

A.Change the job type from Python to Spark.
B.Use multiple Glue jobs triggered sequentially.
C.Increase the RDS instance size to improve read throughput.
D.Use AWS Glue Spark job with 100 workers.
AnswerD

More workers enable parallelism, reducing runtime.

Why this answer

AWS Glue Spark jobs can parallelize data processing across multiple workers, dramatically reducing runtime. With 100 workers, the job can process the 500 GB dataset in parallel, achieving sub-1-hour runtime while keeping costs relatively low since Glue charges per DPU-second and the total DPU-seconds may be similar to the original 6-hour single-DPU job.

Exam trap

The trap here is that candidates might think increasing parallelism (Option D) is too expensive, but Glue's pay-per-DPU-second model means a job with 100 workers running for 1 hour costs roughly the same as 1 worker running for 100 hours, so the total cost is similar, not significantly higher.

How to eliminate wrong answers

Option A is wrong because changing from Python to Spark alone does not add parallelism; the job still runs on a single DPU unless the number of workers is increased. Option B is wrong because running multiple Glue jobs sequentially would increase total runtime, not reduce it, as each job would still process data serially. Option C is wrong because the bottleneck is Glue's processing capacity, not RDS read throughput; increasing RDS instance size would not significantly reduce Glue job runtime since the job already reads 500 GB over 6 hours, and the read rate is not the limiting factor.

997
MCQhard

A data engineer is configuring an Amazon S3 bucket to store sensitive financial data. The company requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption key be automatically rotated every year. The engineer creates a KMS customer managed key and enables automatic rotation. When uploading objects using the AWS CLI, the engineer uses the --sse aws:kms parameter but does not specify a key ID. What is the result of this configuration?

A.The objects are encrypted with the customer managed key because it is the only KMS key in the account.
B.The objects are encrypted with the customer managed key, and automatic rotation applies because the key is the default KMS key for the account.
C.The objects are encrypted with the AWS managed key for S3 (aws/s3), not the customer managed key.
D.The upload fails because a KMS key ID is required when using --sse aws:kms.
AnswerC

When you specify --sse aws:kms without a key ID, S3 uses the AWS managed key for S3 (aws/s3) by default. This key is managed by AWS and does not support automatic rotation configuration by the customer. To use the customer managed key, the engineer must specify its key ARN or alias with --sse-kms-key-id.

Why this answer

To use a specific KMS customer managed key for S3 default encryption or per-object encryption, you must explicitly provide the key ARN or alias. Omitting the key ID causes S3 to fall back to the AWS managed key (aws/s3), which does not meet the requirement of using a customer managed key with annual rotation.

Exam trap

The trap here is assuming that specifying --sse aws:kms automatically uses a customer managed key, when in fact it defaults to the AWS managed key unless a key ID is provided.

998
Multi-Selecthard

A company uses AWS DMS to replicate data from an Amazon RDS for MySQL database to Amazon S3. Which TWO configurations are required to enable continuous change data capture (CDC) from MySQL?

Select 2 answers
A.Ensure the S3 bucket is in the same AWS Region as the source database
B.Grant REPLICATION CLIENT and REPLICATION SLAVE privileges to the DMS user
C.Enable binary logging (binlog) on the MySQL source database
D.Enable versioning on the target S3 bucket
E.Configure the MySQL source to be Multi-AZ
AnswersB, C

DMS requires REPLICATION CLIENT to read binlog metadata and REPLICATION SLAVE to stream binlog events as a replica. Granting both to the DMS user satisfies the stem's continuous CDC requirement, since MySQL denies replication access without these specific global privileges.

Why this answer

Option B is correct because AWS DMS requires the MySQL endpoint user to have REPLICATION CLIENT (to read binlog metadata/status) and REPLICATION SLAVE (to read the binary log events) privileges so it can capture ongoing changes for CDC. Option C is correct because continuous change data capture depends on the MySQL source having binary logging enabled (log_bin=ON) with a suitable binlog_format (e.g., ROW) so DMS can read committed changes from the binlog. Option A is not required, since DMS can replicate to an S3 bucket in a different Region than the source database.

Option D is not required, as S3 bucket versioning is unrelated to DMS CDC and is not needed for change capture. Option E is not required, because Multi-AZ is a high-availability deployment choice for the source and does not enable or affect DMS change data capture.

Exam trap

DEA-C01 often tests the specific prerequisites for MySQL CDC with DMS. Candidates might think that Multi-AZ or S3 versioning is required, but the key requirements are binlog and replication privileges. It's also common to confuse REPLICATION CLIENT with other privileges like SELECT.

999
Multi-Selecthard

Which THREE factors should be considered when choosing a partition key for an Amazon DynamoDB table?

Select 3 answers
A.The partition key should be chosen to maximize the size of items in each partition.
B.If the table has a write-heavy workload, the partition key should distribute writes evenly.
C.The partition key should align with the most common query access pattern.
D.The partition key should be chosen to minimize read capacity unit consumption.
E.The partition key should have high cardinality to distribute data evenly.
AnswersB, C, E

Even write distribution across partitions prevents hot partitions, which throttle throughput when write volume is high. DynamoDB hashes the partition key to place items, so a high-cardinality key spreads writes evenly. This directly satisfies the stem's write-heavy workload constraint, sustaining provisioned capacity without request throttling.

Why this answer

Option B is correct because a write-heavy workload requires the partition key to spread write traffic across many partitions; otherwise a hot partition throttles throughput, since each partition supports a limited write capacity (up to 1,000 WCU per partition). Option C is correct because DynamoDB retrieves items by partition key, so aligning the key with the most common query access pattern enables efficient Query operations instead of expensive Scan operations. Option E is correct because high cardinality produces many distinct partition key values, which distributes items and traffic evenly across partitions and avoids hot partitions.

Option A is not a valid factor because item size does not determine partition key choice; DynamoDB partitions data by key value, and large items only consume more capacity, not improve partitioning. Option D is not a valid factor because RCU consumption is driven by item size and consistency model, not by selecting a partition key to minimize reads.

Exam trap

The trap here is that candidates may think maximizing item size (Option A) or minimizing RCU consumption (Option D) are primary factors, when in fact even distribution and access pattern alignment are the critical design principles for DynamoDB partition keys.

1000
MCQhard

A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to an S3 bucket. The data is delivered in 5-minute intervals. However, the engineer notices that the data in S3 is often delayed by up to 30 minutes. Which configuration change would most likely reduce the delay?

A.Decrease the 'Buffer interval' from 300 seconds to 60 seconds.
B.Enable compression (GZIP) on the Firehose delivery stream.
C.Increase the 'Buffer size' from 5 MB to 50 MB.
D.Enable 'Dynamic partitioning' on the Firehose stream.
AnswerA

Firehose buffers incoming records until the buffer interval elapses or the buffer size fills, so a 300-second interval can compound with delivery and retry latency. Lowering it to 60 seconds flushes data to S3 more frequently, directly cutting the observed delay.

Why this answer

The buffer interval determines the maximum time Firehose will wait before delivering data, regardless of buffer size. Decreasing it from 300 to 60 seconds forces more frequent deliveries, reducing the delay. Option B (compression) reduces data size, which could slow buffer filling and potentially increase delay if the buffer size trigger is not met.

Option C (increasing buffer size) would cause Firehose to wait longer for the buffer to fill, increasing delay. Option D (dynamic partitioning) affects data organization, not delivery frequency, so it does not reduce delay.

1001
MCQmedium

A data engineer manages an Amazon S3 data lake with millions of small JSON files ingested continuously. Amazon Athena queries over this data are slow and expensive because each query scans many small objects. The engineer wants to improve query performance and reduce cost without changing the data content. Which solution should the engineer implement?

A.Use AWS Glue ETL to compact the small files into larger Parquet files partitioned by common query filters.
B.Increase the Athena query result reuse cache TTL to 7 days.
C.Convert the S3 bucket to S3 Intelligent-Tiering to improve read throughput.
D.Enable S3 Transfer Acceleration on the bucket to speed up Athena query reads.
AnswerA

Compacting small JSON files into larger Parquet files reduces the number of objects Athena must list and open, and Parquet's columnar format allows Athena to scan only needed columns. Partitioning by common filters further reduces data scanned. This directly addresses both performance and cost without altering the underlying data semantics, making it the most effective solution for this scenario.

Why this answer

The scenario describes slow and costly Athena queries caused by many small JSON files. Compacting into larger Parquet files and partitioning by common filters reduces the number of objects scanned and leverages columnar storage to scan less data. This directly improves performance and lowers cost without changing data content, making it the correct solution.

Exam trap

The trap here is assuming that S3 storage-class or transfer features can improve Athena query performance, when the real issue is file format and object count.

1002
MCQmedium

A data engineer must mask the last four digits of a credit card column in an Amazon Redshift table so that analysts in a specific role see masked values while a fraud team sees the full values. The engineer wants a solution that applies to all queries without modifying each analyst's SQL. Which approach should the engineer use?

A.Create a view that applies a masking expression to the credit card column and grant the analyst role access only to the view.
B.Encrypt the credit card column with AWS KMS and grant the analyst role decrypt permissions on a different key than the fraud team.
C.Use row-level security to restrict the analyst role to rows where the credit card column is not null.
D.Create a dynamic data masking policy with the credit card column, then attach it to the analyst role so masking applies automatically at query time.
AnswerD

Redshift dynamic data masking policies are attached to roles and applied automatically whenever a user in that role queries the column, so no SQL changes are required. The fraud team, not attached to the masking policy, continues to see full values. This satisfies both the blanket masking requirement and role-based visibility in a single column-level control.

Why this answer

Redshift dynamic data masking attaches a masking policy to a role, so any query by that role against the tagged column returns masked values automatically, with no SQL rewriting. Users in roles without the policy see the original data. This provides column-level, role-based protection that satisfies both teams' visibility needs.

Exam trap

The trap here is choosing row-level security or a masking view, when only dynamic data masking changes column values transparently per role without altering queries.

1003
MCQmedium

A data engineer is configuring an AWS Glue crawler against an Amazon S3 path that contains CSV files with inconsistent column counts across files. The crawler keeps creating multiple tables for the same data and the engineer wants a single table with a merged schema. Which crawler configuration should the engineer change?

A.Increase the crawler's 'Maximum number of tables' limit in the crawler configuration.
B.Enable the crawler's 'Update all new and existing partitions with metadata from the table' option.
C.Set the crawler's 'Create a single schema for each S3 path' option to true.
D.Change the crawler's classifier to a custom classifier that forces all columns to string type.
AnswerC

This option, sometimes called 'CombineCompatibleSchemas', tells the crawler to merge compatible schemas that share the same S3 path into one table. When CSV files have varying column counts but compatible types, enabling this setting produces a single table with the union of columns instead of multiple tables, which is exactly what the engineer needs.

Why this answer

The 'Create a single schema for each S3 path' crawler setting merges compatible schemas found under the same S3 path into one table. For CSV files with inconsistent column counts, enabling this option produces a single table with the union of compatible columns, eliminating the multiple tables the crawler was creating.

Exam trap

The trap here is assuming that a custom classifier or partition metadata option controls table grouping, when only the single-schema-per-S3-path setting merges compatible schemas into one table.

1004
MCQeasy

A company needs to transform JSON data from an S3 bucket into a structured format for Amazon Redshift. The transformation should be done serverlessly. Which service should be used?

A.AWS Glue
B.Amazon EMR
C.Amazon Athena
D.AWS Lambda
AnswerA

AWS Glue is serverless, so no cluster management is needed, and its ETL jobs transform JSON from S3 into a structured schema that loads directly into Amazon Redshift, satisfying the serverless transformation constraint in the stem.

Why this answer

AWS Glue is the correct choice because it is a fully managed, serverless ETL service designed specifically for transforming and preparing data for analytics, including converting JSON to structured formats like Parquet or ORC for Amazon Redshift. It can crawl the S3 source, infer schemas, and run Spark-based transformation jobs without provisioning any infrastructure, aligning perfectly with the serverless requirement.

Exam trap

The trap here is that candidates often confuse Amazon Athena's serverless SQL querying capability with ETL transformation, but Athena cannot transform or write data into a different format for Redshift—it only reads and queries data in place.

How to eliminate wrong answers

Option B (Amazon EMR) is wrong because it requires provisioning and managing EC2 clusters, which is not serverless; it is a managed Hadoop framework but still involves underlying infrastructure. Option C (Amazon Athena) is wrong because it is a serverless query engine for analyzing data directly in S3 using SQL, not a transformation service for converting JSON to a structured format for Redshift. Option D (AWS Lambda) is wrong because it is designed for short-running, event-driven functions (max 15-minute execution time) and is not suitable for large-scale ETL transformations on big datasets, which typically require longer-running jobs.

1005
MCQhard

A data engineer maintains an AWS Glue job that incrementally processes new files in Amazon S3 using job bookmarks. After a schema change in the source data added a new column, the engineer updated the Glue Data Catalog table. Subsequent job runs still process only previously seen files and ignore newly arrived objects. The engineer verifies that new files exist in the prefix and that the bookmark state was not reset. Which factor most likely explains why new files are being skipped?

A.Amazon S3 eventual consistency delays object visibility, so bookmarks never see files written within the same hour.
B.Job bookmarks only support Amazon S3 sources when the files are in Parquet format, so JSON files are silently skipped.
C.The Glue job's IAM role lacks s3:ListBucket permission, so the job cannot detect newly arrived objects.
D.Job bookmarks use the source path and a transformation context key; changing the job script or the catalog table can invalidate the bookmark state.
AnswerD

Glue job bookmarks track processed data using a key derived from the source path, transformation context, and related job parameters. Editing the script or altering the catalog table can change that context, causing the bookmark to behave unexpectedly and skip new files. Resetting or recreating the bookmark after such changes restores correct incremental processing.

Why this answer

Job bookmarks persist state keyed to the source path, transformation context, and job parameters. Modifying the job script or the referenced Data Catalog table alters that context, which can cause the bookmark to skip newly arrived objects despite their presence. Resetting the bookmark or recreating the job restores expected incremental behavior.

Format, IAM, and S3 consistency are not the cause here.

Exam trap

The trap here is assuming bookmarks only look at file timestamps, when they actually key on transformation context that script or catalog edits can invalidate.

1006
MCQmedium

A financial services company uses AWS Glue ETL jobs to process sensitive customer data stored in Amazon S3. The data is encrypted at rest with SSE-KMS using a customer-managed key. Recently, the security team discovered that the Glue job's IAM role has an overly permissive policy that allows the 'kms:Decrypt' action for all KMS keys in the account. The company wants to follow the principle of least privilege. The Glue job runs on a schedule and reads from a specific S3 bucket. The security team needs to update the IAM policy to restrict KMS decryption to only the specific key used for that bucket. What should they do?

A.Update the policy to allow 'kms:Decrypt' with a resource of 'arn:aws:kms:us-east-1:123456789012:key/*' to cover all keys in the account.
B.Update the policy to allow 'kms:Decrypt' with a resource of '*' to ensure the job can always decrypt data.
C.Update the policy to allow 'kms:Decrypt' only for the specific KMS key ARN used by the S3 bucket containing the customer data.
D.Remove the 'kms:Decrypt' action from the policy and rely on S3 bucket policies to grant decryption permissions.
AnswerC

Restricting the Resource element to the specific KMS key ARN satisfies the least-privilege constraint, since AWS evaluates the Resource field against the key's ARN during the kms:Decrypt authorisation call. The Glue job's IAM role then decrypts only data encrypted under that customer-managed key, eliminating access to every other key in the account.

Why this answer

To follow least privilege, the IAM role for the Glue job should only have access to decrypt using the specific KMS key that encrypts the S3 bucket containing the customer data. This is done by allowing 'kms:Decrypt' with a resource set to the exact ARN of that key, not a wildcard or all keys. Option A is incorrect because using a wildcard in the key ARN (key/*) still grants access to all keys under that key hierarchy, which is overly permissive.

Option B is incorrect because allowing 'kms:Decrypt' with resource '*' would grant access to all keys in the account, violating least privilege. Option D is incorrect because removing 'kms:Decrypt' from the IAM policy would prevent the Glue job from decrypting the data; the job's IAM role needs the permission, and relying solely on S3 bucket policies cannot grant decryption permissions cross-account or for IAM roles.

1007
MCQmedium

A data engineer needs to allow an IAM user to rotate the secret in AWS Secrets Manager for an RDS database. Which IAM action should be included in the policy?

A.secretsmanager:RotateSecret
B.secretsmanager:PutSecretValue
C.secretsmanager:UpdateSecret
D.secretsmanager:GetSecretValue
AnswerA

The `secretsmanager:RotateSecret` action authorises triggering immediate rotation of a stored secret, satisfying the requirement to rotate the RDS credential. It differs from `PutSecretValue`, which merely writes a new value without invoking the rotation Lambda, and from `UpdateSecret`, which modifies metadata or encryption rather than performing rotation.

Why this answer

The secretsmanager:RotateSecret action allows the user to initiate rotation of a secret. Option A is correct. secretsmanager:GetSecretValue only retrieves the secret value, not rotate it.

1008
MCQeasy

A data engineer needs to store semi-structured JSON data from IoT devices. The data is written frequently and read occasionally. Which AWS service is MOST cost-effective for this use case?

A.Amazon ElastiCache for Redis
B.Amazon DynamoDB
C.Amazon RDS for MySQL
D.Amazon Redshift
AnswerB

DynamoDB is a serverless key-value store billed per request, so frequent writes and occasional reads incur no idle capacity cost, unlike always-on provisioned databases. It natively stores JSON documents, satisfying the semi-structured IoT payload requirement at the lowest cost.

Why this answer

Amazon DynamoDB is the most cost-effective choice because it is a fully managed NoSQL key-value and document database that natively supports semi-structured JSON data, offers single-digit millisecond latency for frequent writes, and provides a pay-per-request pricing model ideal for workloads with occasional reads. Its on-demand capacity mode automatically scales to handle high write throughput without provisioning, making it cheaper than provisioned alternatives for spiky or unpredictable IoT ingestion patterns.

Exam trap

The trap here is that candidates often choose Amazon ElastiCache for Redis due to its speed and JSON module support, but they overlook that it is not designed for durable, cost-effective long-term storage of semi-structured data, and DynamoDB's native JSON support and pay-per-request pricing make it the more economical choice for this specific write-frequent, read-occasional pattern.

How to eliminate wrong answers

Option A is wrong because Amazon ElastiCache for Redis is an in-memory cache designed for sub-millisecond read-heavy workloads and ephemeral data, not for durable storage of semi-structured JSON from IoT devices; it lacks native JSON document storage (though RedisJSON module exists, it adds cost and complexity) and is significantly more expensive per GB than DynamoDB for persistent data. Option C is wrong because Amazon RDS for MySQL is a relational database that requires schema definition, making it inefficient for semi-structured JSON data that varies in fields; it incurs higher costs due to provisioned IOPS and storage, and its write performance is limited by the underlying instance size and transaction overhead. Option D is wrong because Amazon Redshift is a columnar data warehouse optimized for complex analytical queries on large datasets, not for high-frequency writes from IoT devices; its minimum cost is high (starts at ~$0.25/hour for dc2.large), and it is overkill for occasional reads of semi-structured JSON, leading to wasted expenditure.

1009
MCQeasy

A data engineer needs to ensure that an Amazon S3 bucket containing sensitive customer data is encrypted at rest. Which AWS service can be used to manage the encryption keys?

A.AWS Certificate Manager
B.AWS Secrets Manager
C.AWS CloudHSM
D.AWS Key Management Service (KMS)
AnswerD

AWS KMS centrally creates, rotates and controls the customer-managed keys used for S3 server-side encryption (SSE-KMS), satisfying the encryption-at-rest requirement. KMS also provides audit trails of key usage through CloudTrail, giving the key management capability the scenario demands.

Why this answer

AWS Key Management Service (KMS) is the managed service for creating, rotating, and controlling the keys used to encrypt data at rest in S3 (SSE-KMS). It integrates natively with S3 so that objects are encrypted with a KMS customer master key, and access to the key is governed by IAM and key policies. This is the standard answer for managing S3 encryption keys.

Exam trap

DEA-C01 often tests whether candidates confuse services that manage different secret types — ACM handles TLS certificates, Secrets Manager handles credentials, and KMS handles encryption keys, so candidates who pick ACM or Secrets Manager for S3 key management fall into the trap.

How to eliminate wrong answers

Option A is wrong because AWS Certificate Manager manages TLS/SSL certificates for encryption in transit, not data-at-rest encryption keys. Option B is wrong because Secrets Manager stores and rotates secrets such as database credentials and API keys, not encryption keys for S3 objects. Option C is wrong because CloudHSM is a dedicated hardware security module for custom key management and is not the managed service used for standard S3 SSE-KMS encryption; it is used when you need single-tenant HSM control.

1010
MCQmedium

A data engineer notices that an AWS Glue ETL job is failing with an OutOfMemory error when processing a large dataset. The job uses a Standard worker type. Which action is MOST effective to resolve this issue without changing the job script?

A.Increase the number of workers
B.Switch to G.1X worker type
C.Change to G.2X worker type
D.Increase the number of DPUs per worker
AnswerC

Changing to G.2X allocates 2 DPU (16 GB) per worker, directly increasing memory and resolving the OutOfMemory error most effectively.

Why this answer

The most effective action is to switch to the G.2X worker type (Option C). AWS Glue worker types determine the vCPU, memory, and disk allocated per worker. Standard and G.1X both provide 4 vCPU and 16 GB memory per worker (1 DPU); G.1X does not increase memory over Standard, only disk.

G.2X doubles resources to 8 vCPU and 32 GB memory (2 DPU), directly resolving the OutOfMemory error without changing the script. Option A increases the number of workers but not per-worker memory. Option B (G.1X) does not increase memory, so it is not effective.

Option D is invalid because DPUs per worker is not a configurable parameter; it is defined by the worker type.

Exam trap

Candidates may think increasing the number of workers adds more memory, but it only increases parallelism. The memory per worker remains the same. Changing worker type (e.g., to G.2X) is the direct fix.

1011
MCQmedium

A data engineer is using AWS Glue to process data stored in Amazon S3. The engineer needs to ensure that the AWS Glue job can access the S3 bucket securely without hardcoding credentials. Which approach should the engineer use?

A.Embed the AWS access key and secret key directly in the Glue job script.
B.Create an IAM role with the necessary S3 permissions and attach it to the AWS Glue job.
C.Store AWS credentials in AWS Secrets Manager and retrieve them within the Glue job script.
D.Use an Amazon S3 bucket policy that allows public read access to the data.
AnswerB

AWS Glue jobs assume an IAM role that grants permissions to access AWS resources. By attaching a role with the appropriate S3 permissions, the job can securely access the bucket without embedding credentials. This is the recommended best practice for AWS services. It also allows for fine-grained access control and auditing.

Why this answer

Attaching an IAM role with the necessary S3 permissions to the AWS Glue job is the secure and recommended way to grant access without hardcoding credentials. IAM roles provide temporary credentials and follow the principle of least privilege. The other options either introduce security risks or unnecessary complexity.

Exam trap

The trap here is considering Secrets Manager or hardcoded credentials when IAM roles are the native, secure method for service-to-service authentication.

1012
MCQhard

A data engineer is designing a multi-region disaster recovery solution for Amazon RDS for PostgreSQL. The primary region must have a standby in a different Availability Zone, and the secondary region must have a readable replica that can be promoted in case of failure. Which configuration meets these requirements?

A.Use a single-AZ primary and enable automatic backups
B.Enable Multi-AZ in the primary region and create a cross-region read replica
C.Use a single-AZ primary and create a cross-region read replica
D.Enable Multi-AZ in both primary and secondary regions
AnswerB

Multi-AZ maintains a synchronous standby in a separate Availability Zone for automatic failover, while a cross-region read replica provides a readable copy in the secondary region that can be promoted. Together they satisfy both the in-region standby and cross-region readable replica constraints.

Why this answer

It meets both requirements: Multi-AZ in the primary region provides a synchronous standby in a different Availability Zone for high availability, and a cross-region read replica in the secondary region provides an asynchronous, readable copy that can be promoted to a standalone primary during a regional failure. This combination ensures both intra-region fault tolerance and inter-region disaster recovery.

Exam trap

The trap here is that candidates often confuse Multi-AZ (synchronous, for high availability within a region) with cross-region read replicas (asynchronous, for disaster recovery), and may incorrectly assume that Multi-AZ alone provides cross-region failover or that a single-AZ primary with a read replica satisfies the intra-region standby requirement.

How to eliminate wrong answers

Option A is wrong because a single-AZ primary with automatic backups does not provide a standby in a different Availability Zone, nor does it create a readable replica in a secondary region; backups are for point-in-time recovery, not for immediate failover or read scaling. Option C is wrong because a single-AZ primary lacks the required standby in a different Availability Zone within the primary region; the cross-region read replica only addresses the secondary region requirement. Option D is wrong because enabling Multi-AZ in both regions does not create a cross-region read replica; Multi-AZ in the secondary region provides a standby within that region but does not establish a readable replica that can be promoted from the primary region.

1013
MCQeasy

A company uses Amazon S3 to store raw data and AWS Glue to run ETL jobs. The data is partitioned by date in the format 'year=YYYY/month=MM/day=DD'. A new data source started sending data with a different date format 'YYYY-MM-DD'. The Glue crawler is configured to create a single table for the entire bucket. The crawler runs daily, but it is not detecting the new partitions from the new data source. The existing partitions are in the format 'year=2024/month=05/day=10', while the new data is stored as '2024-05-10/' without the key-value structure. How should the engineer modify the data pipeline to include the new data?

A.Run the crawler with the 'Create partition indexes' option enabled.
B.Configure the crawler to add a custom classifier for date formats.
C.Modify the new data source to store data in the same Hive-style partition format as the existing data.
D.Convert the new data to Parquet format.
AnswerC

Hive-style partitioning requires the key=value directory structure that Glue's crawler parses to infer partition columns. The new source's flat 'YYYY-MM-DD/' folders lack that structure, so the crawler cannot register them as partitions. Conforming the new data to 'year=YYYY/month=MM/day=DD' satisfies the single-table constraint and restores daily partition detection.

Why this answer

The AWS Glue crawler relies on Hive-style partitioning (e.g., 'year=YYYY/month=MM/day=DD') to automatically detect and create partitions. The new data source uses a different format ('YYYY-MM-DD/') that does not follow the key-value structure, so the crawler cannot recognize it as partitions. The engineer must modify the new data source to store data in the same Hive-style partition format as the existing data.

Exam trap

The trap is thinking that a crawler configuration change (like custom classifiers or partition indexes) can handle non-standard partition formats; actually, the data must be stored in a compatible format for automatic detection.

How to eliminate wrong answers

Option A is wrong because partition indexes improve query performance but do not help the crawler detect partitions with non-Hive-style formats. Option B is wrong because custom classifiers are for parsing data formats (e.g., CSV, JSON), not for interpreting partition structures. Option D is wrong because converting to Parquet changes the file format, not the partition structure, and does not address the crawler's inability to detect the new partitions.

1014
Multi-Selectmedium

A company is using AWS Lake Formation to manage permissions on a data lake. Which of the following are valid ways to grant access to a user or role? (Choose THREE.)

Select 3 answers
A.Grant permissions to a SAML or SCIM group
B.Grant permissions using tag-based access control (LF-Tags)
C.Grant permissions to an IAM user or role
D.Grant permissions to an AWS Organizations unit
E.Grant permissions via an S3 bucket policy
AnswersA, B, C

Lake Formation integrates with SAML and SCIM identity providers, so permissions granted to an external group propagate to its members without per-user grants. This satisfies the stem's requirement for valid access-granting methods, alongside IAM principals and tag-based LF-TBAC grants, reducing administrative overhead.

Why this answer

Lake Formation supports granting permissions directly to IAM users and roles (option C), which is the fundamental way principals are identified in AWS. It also supports granting permissions to SAML or SCIM groups (option A), allowing federated identities to inherit Lake Formation permissions through their group membership. Additionally, Lake Formation supports tag-based access control using LF-Tags (option B), where permissions are granted on tags and then associated with resources and principals.

Option D is incorrect because Lake Formation does not grant permissions to AWS Organizations units; it works with IAM principals and federated groups, not OUs directly. Option E is incorrect because S3 bucket policies are an S3-level access mechanism and do not grant Lake Formation permissions, which are managed within Lake Formation's own permission model.

Exam trap

Tag-based access control (LF-Tags) is a valid method in Lake Formation, similar to IAM resource tags, but it is specific to Lake Formation.

1015
MCQmedium

A data engineer is troubleshooting an Amazon Redshift cluster that is not responding to queries. The engineer suspects that the cluster may have been accidentally deleted. Which AWS service should be used to investigate the deletion?

A.AWS Config
B.AWS CloudTrail
C.Amazon CloudWatch Logs
D.AWS Trusted Advisor
AnswerB

AWS CloudTrail records DeleteCluster API calls, capturing the identity, source IP and timestamp of the deletion. This event history lets the engineer confirm whether the cluster was deleted, by whom, and when, which CloudWatch metrics or cluster logs alone cannot establish.

Why this answer

AWS CloudTrail records API activity in an AWS account, including the DeleteCluster API call that would be logged when a Redshift cluster is deleted. By querying CloudTrail event history or a trail's logs, the engineer can identify who deleted the cluster, when, and from which source IP. This makes CloudTrail the correct service for investigating the deletion event itself.

Exam trap

DEA-C01 often tests the confusion between CloudTrail (API audit/activity) and AWS Config (resource configuration history), since both can show that a resource no longer exists but only CloudTrail reveals who made the API call.

How to eliminate wrong answers

Option A is wrong because AWS Config records resource configuration changes and compliance state over time, but it does not capture the API caller identity or the specific delete request details needed to investigate who performed the deletion. Option C is wrong because CloudWatch Logs stores application and service log output, not AWS API audit events — Redshift cluster deletion is an API action, not a log stream event. Option D is wrong because Trusted Advisor provides best-practice checks and recommendations, not an audit trail of API calls or deletion events.

1016
Multi-Selectmedium

A company uses Amazon DynamoDB for a gaming application. The application experiences throttling during peak hours. The table's read and write capacity is provisioned. Which TWO actions can reduce throttling?

Select 2 answers
A.Enable TTL (time to live) on the table to automatically delete old items
B.Enable DynamoDB auto scaling for the table
C.Increase the provisioned read capacity units (RCUs)
D.Implement DynamoDB Accelerator (DAX) to cache read requests
E.Add a DynamoDB Global Table for the table
AnswersB, D

Auto scaling adjusts provisioned capacity based on traffic.

Why this answer

DynamoDB auto scaling (Option B) automatically adjusts the provisioned read and write capacity based on actual traffic patterns, preventing throttling during peak hours without manual intervention. This is the correct action because it dynamically increases capacity when demand spikes and reduces it during low traffic, directly addressing the throttling issue.

Exam trap

The trap here is that candidates often confuse increasing provisioned capacity (Option C) as the only solution, but the exam tests whether you understand that auto scaling (Option B) is the correct managed approach, and that DAX (Option D) can reduce read throttling by caching, making both B and D valid together.

1017
MCQmedium

A company is using Amazon Kinesis Data Streams to ingest real-time clickstream data from a website. The data is consumed by an Amazon Kinesis Data Analytics for Apache Flink application that performs real-time analytics. The Flink application writes its results to an Amazon S3 bucket. The company has noticed that the Flink application is experiencing high checkpoint failure rates, causing delays. The CloudWatch metrics show that the checkpoint size is large and increasing. The data engineer needs to reduce the checkpoint size. Which action should the data engineer take?

A.Decrease the checkpoint interval to reduce the amount of state accumulated.
B.Reduce the parallelism of the Flink application.
C.Increase the state time-to-live (TTL) configuration to retain state longer.
D.Enable incremental checkpointing in the Flink application to only write changes since the last checkpoint.
AnswerD

Incremental checkpointing writes only the state changes since the previous checkpoint rather than the full state, so checkpoint size stops growing with application state. This directly addresses the large and increasing checkpoint size causing the high failure rates.

Why this answer

Enabling incremental checkpointing in the Flink application is the correct action because it only writes the changes since the last checkpoint, significantly reducing checkpoint size. This is especially effective when state is large and growing, as it avoids rewriting the entire state with each checkpoint.

Exam trap

DEA-C01 often tests Flink checkpoint tuning. Candidates may think that reducing checkpoint interval or parallelism helps, but the trap is not knowing that incremental checkpointing is the key to reducing checkpoint size for large state.

How to eliminate wrong answers

Option A is wrong because decreasing the checkpoint interval would increase checkpoint frequency, potentially exacerbating the issue by creating more checkpoints, though each might be smaller; it does not address the root cause of large state size. Option B is wrong because reducing parallelism would decrease throughput and may not reduce checkpoint size; it could even increase per-task state. Option C is wrong because increasing state TTL would retain state longer, making the state larger and worsening the problem.

1018
MCQhard

A data engineer runs an AWS Glue Studio job that reads JSON from Amazon S3 and writes to a partitioned Parquet table. The job currently runs for six hours. Profiling shows that a small number of partitions contain millions of rows while most contain a few hundred. Which change will MOST improve runtime?

A.Increase the number of DPUs allocated to the Glue job from 10 to 100.
B.Apply a salting technique to the skewed keys before the write, or repartition on a composite key.
C.Enable job bookmarks and re-run the job to skip already-processed files.
D.Switch the job's worker type from G.1X to G.2X to get more memory per executor.
AnswerB

Salting appends a random suffix to hot keys so rows spread across many tasks, and repartitioning on a composite key achieves similar balance. Both directly address the uneven task durations causing the long tail. This is the standard remedy for data skew in Spark and therefore in AWS Glue jobs.

Why this answer

Data skew, not compute capacity, is causing the long runtime because a few partitions hold most of the rows. Salting hot keys or repartitioning on a composite key spreads those rows across many tasks so no single task becomes a straggler. Increasing DPUs, changing worker type, or enabling bookmarks does not change how rows are distributed within a run.

Exam trap

The trap here is assuming that more DPUs or bigger workers will fix any slow job, when skew requires redistributing data rather than adding compute.

1019
MCQmedium

A media company stores millions of thumbnail images in an Amazon S3 bucket. Analysts run ad hoc queries against the image metadata, which is kept as JSON objects in the same bucket. Query latency is unpredictable and costs are rising because Athena scans large volumes of JSON for every query. The team wants faster queries and lower scan cost while keeping the data in S3 and queryable with SQL. Which change should the data engineer make?

A.Use AWS Glue to crawl the metadata, convert it to Apache Parquet partitioned by date, and register the table in the Data Catalog for Athena queries.
B.Enable S3 Intelligent-Tiering on the bucket to automatically move infrequently accessed metadata objects to cheaper storage.
C.Move the metadata into an Amazon DynamoDB table and have analysts query it with PartiQL.
D.Increase the Athena workgroup data usage control limit so queries can scan more data without failing.
AnswerA

Converting JSON metadata to columnar Parquet lets Athena read only the referenced columns and benefits from compression, while date partitioning restricts each query to relevant folders. Registering the table in the Data Catalog makes it directly queryable, reducing scanned bytes and cost while keeping the data in S3.

Why this answer

The root cause is that Athena must read entire JSON objects and every partition for each query. Converting metadata to columnar Parquet and partitioning by date lets the engine read only needed columns and folders, cutting scanned bytes and latency. Catalog registration keeps the data in S3 and preserves SQL access, which matches all stated constraints.

Exam trap

The trap here is assuming that cheaper S3 storage tiers or higher query limits reduce Athena scan cost, when the real lever is the data format and partitioning.

1020
MCQmedium

A data engineer is monitoring an Amazon Kinesis Data Stream and notices that the 'WriteProvisionedThroughputExceeded' metric is frequently elevated. The stream has 5 shards and is used by multiple producers. What is the BEST action to resolve this issue?

A.Increase the consumer's processing speed to reduce lag.
B.Increase the number of shards in the Kinesis data stream.
C.Reduce the data retention period of the stream.
D.Implement exponential backoff and retries in the producer applications.
AnswerB

WriteProvisionedThroughputExceeded means producers are exceeding the per-shard ingest limit of 1 MB/s or 1,000 records/s. Adding shards increases total stream capacity by distributing writes across more shards, directly resolving the throttling described in the stem.

Why this answer

WriteProvisionedThroughputExceeded indicates that the write rate exceeds the shards' capacity. Increasing the number of shards increases the total write capacity. Option A is incorrect because increasing the consumer's processing speed does not affect write throttling; it addresses read-side lag.

Option C is incorrect because reducing the retention period does not affect write throughput. Option D is incorrect because implementing exponential backoff and retries in the producer applications helps with transient failures but does not resolve the root cause of insufficient capacity.

1021
MCQeasy

A data engineer needs to ensure that an AWS Glue job has access to an Amazon RDS database in a private subnet. The Glue job will run in a VPC and requires a security group and subnet configuration. Which combination of steps should the engineer take?

A.Use AWS Glue Studio to create a connection to RDS, which automatically configures the necessary VPC settings and security groups.
B.Configure the Glue job with a VPC connection, specify a subnet and security group, and ensure the security group allows outbound traffic to the RDS database's security group.
C.Attach an IAM role to the Glue job with permissions to access RDS, and configure the RDS database to be publicly accessible.
D.Create a VPC endpoint for AWS Glue, configure the Glue job with a VPC connection, and attach a security group that allows outbound traffic to the RDS database.
AnswerB

To run a Glue job in a VPC, you must provide a VPC connection that includes subnet and security group settings. The security group must allow outbound traffic to the RDS database's security group on the database port. This is the correct and minimal configuration for Glue to access RDS in a private subnet.

Why this answer

For an AWS Glue job to access an Amazon RDS database in a private subnet, the job must be configured with a VPC connection that specifies the subnet and security group. The security group must allow outbound traffic to the RDS database's security group on the database port. This ensures network connectivity while maintaining security.

Exam trap

The trap here is thinking that a VPC endpoint for AWS Glue is needed or that making RDS public is acceptable, when in fact a VPC connection with proper security groups is the correct approach.

1022
MCQmedium

A data engineering team is troubleshooting a failing AWS Glue ETL job that processes data from an S3 bucket. The job writes output to another S3 bucket. The job fails with an AccessDenied error when writing to the output bucket. The IAM role used by the job has the following policy attached: {"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":["s3:GetObject","s3:ListBucket"],"Resource":["arn:aws:s3:::input-bucket/*","arn:aws:s3:::input-bucket"]}]}. What is the most likely cause of the failure?

A.The ETL job is processing more than 10 TB of data.
B.The output bucket has a bucket policy that denies access to the IAM role.
C.The IAM role does not have s3:PutObject permission on the output bucket.
D.The IAM role used by the job does not exist.
AnswerC

The attached policy grants only s3:GetObject and s3:ListBucket on the input bucket. Writing objects requires s3:PutObject, which is absent for the output bucket, so S3 returns AccessDenied. Adding s3:PutObject on the output bucket's ARN resolves the failure.

Why this answer

The IAM policy attached to the Glue job role only grants s3:GetObject and s3:ListBucket on the input bucket. There is no s3:PutObject (or any write) permission on the output bucket, so when Glue attempts to write results, S3 returns AccessDenied. The most likely cause is therefore the missing write permission on the output bucket.

Exam trap

The trap here is that candidates focus on the input bucket permissions (which are present) and overlook that write operations require a completely separate s3:PutObject action that is absent from the policy.

How to eliminate wrong answers

Option A is wrong because data volume (10 TB) is not a cause of AccessDenied errors — Glue has no 10 TB per-job limit that produces AccessDenied. Option B is wrong because while a bucket policy could deny access, the question states the policy attached to the role only allows read on the input bucket, so the missing PutObject permission is the primary and most likely cause; a deny in the output bucket policy is not indicated. Option D is wrong because if the IAM role did not exist, the job would fail at role assumption with a different error (e.g., AccessDenied on sts:AssumeRole or InvalidParameterValue), not specifically on writing to the output bucket.

1023
MCQmedium

A data engineer is building an AWS Glue Studio visual ETL job that reads JSON files from Amazon S3, applies a transformation, and writes to Amazon Redshift. During a test run, the job fails with an error indicating that the dynamic frame could not be written because the target table schema does not match the incoming data. The engineer needs to ensure the job automatically reconciles schema differences such as missing columns and data type mismatches during the write. Which action should the engineer take?

A.Set the job bookmarks option to 'Enable' so Glue tracks schema changes between runs and applies them to the target.
B.Change the write node's 'Handling of schema differences' to 'Reconcile schema' so missing columns are added and type mismatches are cast.
C.Enable the 'Auto schema reconciliation' option in the Glue job's Redshift connection parameters.
D.Increase the number of DPUs allocated to the Glue job so the write can complete with a larger executor pool.
AnswerB

In AWS Glue Studio, the Redshift write node exposes a 'Handling of schema differences' property. Setting it to 'Reconcile schema' lets the writer add missing columns and coerce compatible data types to match the target table. This directly addresses the failure described, since the dynamic frame will be aligned to the Redshift table schema before the COPY operation is issued.

Why this answer

The Redshift write node in AWS Glue Studio provides a 'Handling of schema differences' setting. Selecting 'Reconcile schema' lets the writer add missing columns and cast compatible types so the dynamic frame matches the target table. This is the supported mechanism for automatically aligning incoming data with an existing Redshift schema, and it resolves the described failure without manual DDL changes.

Exam trap

The trap here is assuming that Glue connection settings or job bookmarks control runtime schema behavior, when schema reconciliation is a property of the write node itself.

1024
MCQhard

A data engineer is building an AWS Glue ETL job that reads from an AWS Glue Data Catalog table backed by Amazon S3. The job must process only records added since the last successful run to reduce cost and runtime. The source data is partitioned by year, month, and day. Which approach should the engineer use?

A.Use the AWS Glue Data Catalog's update time and a custom script that queries the catalog for recently changed partitions before reading S3.
B.Set the job's maximum concurrency to 1 and rely on the S3 object's last modified timestamp in the script to filter records.
C.Configure the job to use a pushdown predicate that filters on the year, month, and day partition columns using the current date.
D.Enable job bookmarks in the AWS Glue job and use the default transformation_ctx on the source DynamicFrame.
AnswerD

AWS Glue job bookmarks track the state of previously processed data. When enabled and a transformation context is set on the source, the job reads only new or changed data since the last run. For partitioned S3 sources, bookmarks use partition metadata to determine what is new, which directly satisfies the requirement to process only newly added records.

Why this answer

AWS Glue job bookmarks persist state across runs and, when a transformation context is set on the source, cause the job to read only data that has not been processed before. For partitioned S3 sources, bookmarks track new partitions and files, which precisely matches the requirement to process only records added since the last successful run while reducing cost and runtime.

Exam trap

The trap here is confusing a pushdown predicate that filters by current date with true incremental state tracking across job runs.

1025
Multi-Selectmedium

A data engineer is designing an ingestion pipeline that uses AWS Glue to read from an Amazon RDS for PostgreSQL database. The job must read only rows changed since the previous run and must not scan the entire table each night. The source table has a last_updated timestamp column that is updated on every write. (Choose two.)

Select 2 answers
A.Configure the Glue JDBC connection with a pushdown predicate that filters on last_updated greater than a stored high-water mark.
B.Set the Glue job's '--enable-auto-scaling' argument so the job adds workers as the source table grows.
C.Create the DynamicFrame with the 'additional_options' parameter set to {'hashfield': 'id'} to enable hash-based change detection.
D.Enable job bookmarks on the Glue job so it tracks the last processed state of the JDBC source between runs.
E.Use the Glue ResolveChoice transformation with the 'cast:double' option on the timestamp column.
AnswersA, D

A pushdown predicate is sent to the source database as part of the SQL query, so RDS filters rows before returning them to Glue. Filtering on last_updated greater than a persisted high-water mark reads only changed rows, minimizes network transfer, and satisfies the incremental requirement without full-table scans.

Why this answer

Incremental JDBC ingestion in Glue relies on a pushdown predicate that filters on a high-water mark column like last_updated, combined with job bookmarks that persist state between runs. The predicate pushes work to RDS and reduces transfer, while bookmarks ensure the job resumes from the previous position without reprocessing. Transformations and worker-tuning options do not affect which rows are read.

Exam trap

The trap here is assuming that enabling job bookmarks alone is enough, when a pushdown predicate is also required to restrict the SQL read on the source.

1026
MCQmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The security team requires that data be encrypted at rest using a customer-managed AWS KMS key, and that the Glue job be able to decrypt the source data and encrypt the target data. The engineer has already created a KMS key and attached a key policy that allows the Glue service role to use the key for encrypt and decrypt operations. However, when the job runs, it fails with an access denied error related to KMS. What is the most likely cause of the failure?

A.The KMS key is in a different AWS Region than the S3 buckets, causing cross-Region latency and timeouts.
B.The Glue job is using an outdated version of the AWS SDK that does not support KMS encryption.
C.The Glue job's IAM role lacks permissions for kms:Decrypt and kms:GenerateDataKey.
D.The S3 bucket policy does not allow the Glue job's IAM role to perform s3:GetObject and s3:PutObject.
AnswerC

The Glue job assumes an IAM role to access AWS services. Even if the KMS key policy grants access, the IAM role must also have explicit permissions for kms:Decrypt and kms:GenerateDataKey to use the key for reading and writing encrypted data. Without these IAM permissions, the job cannot decrypt source objects or generate data keys for encryption, resulting in access denied.

Why this answer

For a Glue job to use a customer-managed KMS key, both the key policy and the IAM role's identity-based policy must grant the necessary permissions. The key policy alone is insufficient; the IAM role must explicitly allow kms:Decrypt and kms:GenerateDataKey. This dual authorization ensures least privilege.

The error indicates the IAM role lacks these permissions, so adding them resolves the issue.

Exam trap

The trap here is assuming that a permissive KMS key policy is enough, overlooking that the IAM role also needs explicit permissions for KMS actions.

1027
Multi-Selectmedium

A company is using AWS Glue to run ETL jobs that transform data from S3 to Redshift. The jobs are failing intermittently with out-of-memory errors. Which THREE actions can help resolve this issue? (Choose THREE.)

Select 3 answers
A.Increase the number of DPUs allocated to the Glue job
B.Use S3 Select to filter data before reading into the Glue job
C.Use Spark's 'coalesce' function to reduce the number of partitions
D.Optimize the transformation logic to use less memory, for example by filtering early
E.Use a larger worker type, such as G.2X
AnswersA, D, E

More DPUs provide more memory and compute resources.

Why this answer

Increasing the number of DPUs allocated to the Glue job provides more memory and compute resources for the Spark executors, directly addressing out-of-memory errors by allowing larger datasets to be processed without exceeding heap limits. This is a standard scaling approach for memory-intensive ETL workloads in AWS Glue.

Exam trap

The trap here is that candidates often confuse reducing data volume (S3 Select) with increasing memory capacity, or mistakenly believe coalescing partitions always reduces memory usage, when in fact it can concentrate data and exacerbate OOM errors.

1028
MCQhard

A company has a 100 TB dataset stored on-premises in a Hadoop cluster. They want to ingest this data into Amazon S3 for processing with AWS Glue. The company has a limited time window and a slow internet connection. Which strategy is MOST appropriate?

A.Use AWS Snowball Edge to physically ship the data to AWS.
B.Use AWS DataSync over the existing internet connection.
C.Use Amazon S3 Transfer Acceleration to speed up the upload.
D.Use AWS Direct Connect to establish a high-bandwidth connection.
AnswerA

Snowball Edge bypasses the slow internet connection by physically shipping 100 TB, meeting the limited time window. Network transfer of that volume over a slow link would take far too long, so physical shipment is the appropriate strategy.

Why this answer

AWS Snowball Edge is the most appropriate strategy because the dataset is 100 TB, the time window is limited, and the internet connection is slow. Snowball Edge provides a physical storage device that can be shipped to AWS, bypassing network bandwidth constraints entirely. This approach is designed for large-scale data transfers (typically over 10 TB) where network transfer would be impractical or exceed the available time window.

Exam trap

The trap here is that candidates may overestimate the effectiveness of network acceleration techniques (like Transfer Acceleration or Direct Connect) for extremely large datasets, failing to recognize that physical shipping is the only viable option when bandwidth and time are severely constrained.

How to eliminate wrong answers

Option B is wrong because AWS DataSync relies on the existing internet connection, which is slow and would take an excessively long time to transfer 100 TB, likely exceeding the limited time window. Option C is wrong because Amazon S3 Transfer Acceleration uses edge locations and optimized network paths, but it still depends on the underlying internet connection speed; a slow connection will remain a bottleneck, and it is not designed for petabyte-scale offline transfers. Option D is wrong because AWS Direct Connect requires establishing a dedicated network connection, which involves significant lead time for setup and does not solve the immediate problem of a slow internet connection; it also still transfers data over a network, which for 100 TB would be time-consuming even at high bandwidth.

1029
MCQmedium

A company runs an Amazon EMR cluster with Spark jobs that process data from Amazon S3. The data engineer receives an alert that one of the Spark jobs failed with an OutOfMemoryError. The job processes large files and uses the default Spark configurations. Which configuration change is MOST likely to resolve the issue?

A.Increase the spark.executor.memory configuration.
B.Increase the number of executors.
C.Disable dynamic resource allocation.
D.Decrease the number of cores per executor.
AnswerA

Spark executors hold partitions in heap, so an OutOfMemoryError during processing of large files signals insufficient executor heap. Raising spark.executor.memory directly expands that heap, letting each executor hold its partition data without spilling or failing, which addresses the default-configuration constraint in the stem.

Why this answer

An OutOfMemoryError in Spark typically occurs when an executor's JVM heap is exhausted during processing of large partitions or files. Increasing spark.executor.memory directly expands the heap available to each executor, allowing it to hold more data in memory and resolve the OOM condition.

Exam trap

The trap is assuming that adding more executors (horizontal scaling) fixes a memory error, when OOM is a per-executor heap problem that requires increasing executor memory or fixing data skew.

How to eliminate wrong answers

Option B is wrong because adding more executors increases parallelism but does not increase the memory available to any single executor, so an executor processing a large partition can still OOM. Option C is wrong because disabling dynamic resource allocation removes the ability to scale executors up or down and does not address per-executor memory limits. Option D is wrong because decreasing cores per executor reduces concurrent tasks per executor, which can help in some cases but is not the most direct fix for a memory error; increasing executor memory is the canonical remedy.

1030
MCQmedium

A company uses AWS Glue to process data from multiple sources. The data is stored in an Amazon S3 data lake. The company needs to transform the data using a custom Python library that is not available in the default Glue environment. What is the MOST efficient way to make this library available to the Glue jobs?

A.Manually install the library on each node in the Glue cluster by editing the bootstrap script.
B.Upload the library as a .whl file to Amazon S3 and reference it in the Glue job's --additional-python-modules parameter.
C.Create a custom Docker image with the library and use it in AWS Glue for Ray.
D.Use a shell command in the Glue job script to run 'pip install <library>' before the job runs.
AnswerB

Glue's --additional-python-modules parameter installs wheels from S3 into the job's Python environment at runtime, satisfying the need for a custom library absent from the default Glue image without building a custom connector or repackaging the job.

Why this answer

AWS Glue supports adding custom Python libraries by uploading a .whl file to Amazon S3 and referencing it via the `--additional-python-modules` job parameter. This method is the most efficient as it requires no manual node configuration, no custom Docker images, and no runtime pip installs, ensuring the library is automatically distributed to all worker nodes before the job executes.

Exam trap

The trap here is that candidates may think running 'pip install' directly in the script (Option D) is acceptable, but AWS explicitly recommends using the `--additional-python-modules` parameter for efficiency and reliability, as runtime pip installs can fail due to network timeouts or missing build dependencies.

How to eliminate wrong answers

Option A is wrong because manually editing bootstrap scripts to install the library on each node is inefficient, error-prone, and not scalable; Glue manages cluster lifecycle automatically, so manual node-level modifications are not recommended and can be lost on auto-scaling events. Option C is wrong because AWS Glue for Ray is a specific runtime for distributed Python and Ray-based workloads, not a general-purpose Glue ETL job; using a custom Docker image for Ray adds unnecessary complexity and is not the standard approach for standard Glue ETL jobs. Option D is wrong because running 'pip install' inside the Glue job script is inefficient, adds runtime overhead, may fail due to network restrictions or permissions, and is not the intended way to manage dependencies in Glue; the library must be pre-packaged and referenced via the job parameters.

1031
MCQeasy

A data engineering team needs to transform CSV files stored in Amazon S3 into Parquet format using AWS Glue. The files are partitioned by date and are updated hourly. Which AWS Glue feature should be used to automatically detect the schema and partition structure?

A.AWS Glue Crawler
B.AWS Glue DataBrew
C.AWS Lake Formation
D.Amazon Athena
AnswerA

AWS Glue Crawler scans the S3 data, infers the CSV schema, and registers partition structure in the Glue Data Catalog. Scheduled hourly, it keeps metadata current as new date partitions arrive, enabling the ETL job to read and convert files to Parquet.

Why this answer

AWS Glue Crawler is the correct choice because it automatically scans data in S3, infers the schema (including data types), and detects the partition structure (e.g., date-based partitions like year/month/day) by examining the folder hierarchy. It then populates the AWS Glue Data Catalog with metadata, enabling ETL jobs to read the data without manual schema definition.

Exam trap

AWS often tests the distinction between tools that discover metadata (Crawler) versus tools that consume or transform data (Athena, DataBrew), leading candidates to pick Athena because it can query partitioned data, but it cannot automatically detect the partition structure without a pre-existing catalog.

How to eliminate wrong answers

Option B (AWS Glue DataBrew) is wrong because it is a visual data preparation tool for cleaning and normalizing data, not for automatic schema or partition detection. Option C (AWS Lake Formation) is wrong because it provides centralized security and governance for data lakes, but it does not perform schema discovery or partition detection itself. Option D (Amazon Athena) is wrong because it is a query engine that can read data from the Glue Data Catalog, but it does not automatically detect schemas or partitions; it relies on existing catalog metadata.

1032
MCQmedium

A data engineer is deploying an Amazon Redshift cluster that must be accessible only from within a private VPC and must not have a public IP address. The cluster will be queried by an Amazon EMR cluster in the same VPC and by on-premises BI tools over a VPN connection. Which configuration should the engineer choose?

A.Launch the Redshift cluster in a private subnet group and configure an AWS Site-to-Site VPN connection to the VPC.
B.Launch the Redshift cluster with the publicly accessible setting enabled and attach it to a public subnet group.
C.Launch the Redshift cluster with the publicly accessible setting disabled and attach it to a public subnet group.
D.Launch the Redshift cluster with the publicly accessible setting disabled and attach it to a private subnet group.
AnswerD

Disabling the publicly accessible setting prevents the cluster from receiving a public IP address. Placing it in a private subnet group ensures that only resources within the VPC, such as the EMR cluster, and connected networks over VPN can reach it. This meets the requirement for private-only access without exposing the cluster to the internet.

Why this answer

To ensure a Redshift cluster has no public IP and is reachable only privately, the engineer must disable the publicly accessible setting and use a private subnet group. This prevents internet exposure while allowing access from within the VPC and over VPN. The other options either enable public access or do not explicitly disable it, failing the requirement.

Exam trap

The trap here is assuming that placing a cluster in a private subnet automatically disables public accessibility, when the publicly accessible setting must be explicitly turned off to avoid a public IP.

1033
MCQmedium

A data engineer is configuring an Amazon S3 bucket that stores sensitive customer records for analytics. The security team requires that all data be encrypted at rest with keys that are rotated automatically every year and that access be auditable per key. The engineer must minimize operational overhead. Which encryption configuration should be used?

A.SSE-S3 with bucket key enabled
B.SSE-C with customer-provided keys stored in AWS Secrets Manager
C.SSE-KMS with a customer managed key with automatic rotation enabled
D.SSE-KMS with an AWS managed key (aws/s3)
AnswerC

A customer managed KMS key supports enabling automatic key rotation on an annual schedule, and every use of the key is recorded in AWS CloudTrail for auditing. Object encryption at rest is enforced through SSE-KMS, and the data engineer retains control of the key policy. This satisfies the rotation, auditability, and low operational overhead requirements without managing raw key material.

Why this answer

The requirement combines automatic annual key rotation, per-key auditability, and minimal operational overhead. A KMS customer managed key with automatic rotation enabled provides all three: S3 encrypts objects server-side, CloudTrail records key usage, and rotation occurs yearly without manual intervention. AWS managed keys rotate on a different schedule and offer less policy control, while SSE-S3 and SSE-C lack the needed audit or rotation behavior.

Exam trap

The trap here is assuming that SSE-S3 or an AWS managed KMS key provides configurable annual rotation and per-key auditing, when only a customer managed KMS key with rotation enabled meets both requirements.

1034
MCQhard

A company uses Amazon DynamoDB to store user session data. The table has a partition key of user_id and a sort key of session_start. The workload is read-heavy and eventually consistent reads are acceptable. The table is provisioned with 1000 RCUs and 500 WCUs. During peak hours, the application experiences throttling on read operations, but CloudWatch shows that the consumed read capacity is well below the provisioned amount. What is the most likely cause of the throttling?

A.The table's sort key is causing uneven data distribution across partitions, leading to throttling.
B.The application is using strongly consistent reads, which consume twice the read capacity units and cause throttling.
C.The table has a hot partition because user_id values are not evenly distributed, causing some partitions to exceed their read capacity limits.
D.The table's provisioned read capacity is too low for the workload, and the engineer should increase it to resolve throttling.
AnswerC

DynamoDB partitions have a maximum read capacity of 3000 RCUs per partition. If a few user_id values are accessed much more frequently than others, those items reside on the same partition, causing that partition to throttle even though overall table capacity is underutilized. This is a classic hot partition scenario, and the fix is to distribute the workload more evenly, such as by adding a random suffix to the partition key.

Why this answer

DynamoDB throttling can occur even when overall consumed capacity is below provisioned if a single partition exceeds its per-partition limit. With a partition key like user_id, if some users are much more active, their items concentrate on one partition, causing that partition to throttle. The solution is to distribute the workload more evenly, such as by adding a random suffix to the partition key or using a composite partition key.

Exam trap

The trap here is assuming that throttling always means insufficient provisioned capacity, overlooking per-partition limits that cause hot partitions.

1035
MCQhard

A data pipeline ingests JSON data from an S3 bucket using AWS Glue. The JSON files contain nested structures, and the team wants to flatten them for analysis in Amazon Athena. Which Glue transformation is most appropriate?

A.Filter
B.Join
C.Map
D.Relationalize
AnswerD

Flattens nested JSON into separate tables.

Why this answer

Relationalize is specifically designed to flatten nested JSON into relational tables. Option A (Map) applies a function to each record. Option B (Filter) removes records.

Option C (Join) combines datasets.

1036
Multi-Selecthard

Which TWO are valid approaches to troubleshoot a slow Amazon Redshift query? (Choose two.)

Select 2 answers
A.Check for table locks using STV_LOCKS.
B.Enable encryption on the cluster.
C.Use the EXPLAIN command to review the query execution plan.
D.Run VACUUM on the table.
E.Alter the table to change DISTSTYLE to KEY.
AnswersA, C

STV_LOCKS exposes active table locks that block query progress, directly addressing the slow-query constraint. When another transaction holds a lock on a target table, dependent queries queue and stall; inspecting this view identifies the blocking session so it can be resolved. This is a valid, documented Redshift troubleshooting approach.

Why this answer

Option A is correct because STV_LOCKS is the Amazon Redshift system view that shows current table locks, and a query can appear slow simply because it is blocked waiting on a lock held by another transaction, so checking it is a valid troubleshooting step. Option C is correct because the EXPLAIN command returns the query execution plan, revealing expensive operations such as sequential scans, nested loops, or large data redistributions and broadcasts, which directly identifies why a query is slow. Option B is not a troubleshooting approach for slow queries; enabling cluster encryption is a security configuration and does not affect query performance.

Option D is not a diagnostic step but a maintenance action, and running VACUUM blindly is not a valid troubleshooting approach since it can be costly and should follow diagnosis. Option E is likewise a remediation action that changes the table's distribution style, not a way to investigate the cause of slowness, and it should only be applied after identifying a distribution problem.

Exam trap

DEA-C01 often tests the difference between diagnostic steps (EXPLAIN, STV_LOCKS) and remediation or unrelated actions (VACUUM, encryption, DISTSTYLE change) — candidates who pick remediation steps as 'troubleshooting' fall into the trap.

1037
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. The state machine includes a task that runs an AWS Glue job and then waits for its completion. The engineer notices that the Step Functions execution times out after 15 minutes, even though the Glue job takes about 30 minutes to complete. The Step Functions state machine has a timeout of 1 hour. What is the most likely cause of the timeout?

A.The AWS Glue job is running in a VPC without a NAT gateway, causing network timeouts.
B.The Step Functions task is using the optimized integration for AWS Glue, which has a maximum timeout of 15 minutes.
C.The AWS Glue job is not sending a heartbeat to Step Functions, so the task times out.
D.The Step Functions task has a TimeoutSeconds value set to 900, overriding the state machine timeout.
AnswerD

Correct. In Step Functions, individual tasks can have their own timeout. If the task's TimeoutSeconds is set to 900 (15 minutes), it will fail after that duration regardless of the state machine's overall timeout. This is a common misconfiguration when the task timeout is set lower than the expected job duration. The engineer must increase the task timeout or remove it.

Why this answer

The most likely cause is that the Step Functions task has a TimeoutSeconds parameter set to 900 seconds (15 minutes). This parameter overrides the state machine's timeout for that specific task. To allow the Glue job to run for its full duration, the engineer should increase the task timeout to a value greater than the expected job runtime, or remove it to rely on the state machine's timeout.

Exam trap

The trap here is assuming the state machine's overall timeout is the only timeout that matters, overlooking that individual tasks can have their own shorter timeouts.

1038
MCQmedium

A data engineer manages an AWS Glue job that reads from an Amazon S3 bucket containing PII. The security team requires that the data be encrypted at rest using a customer-managed AWS KMS key, and that the engineer be able to audit key usage. The engineer has already created a KMS key. Which combination of steps should the engineer take to meet these requirements?

A.Use SSE-S3 for the S3 bucket and enable S3 server access logging to track key usage.
B.Enable default encryption on the S3 bucket with SSE-KMS using the AWS managed key aws/s3, and enable AWS CloudTrail data events for S3.
C.Configure the Glue job to use client-side encryption with a customer-managed KMS key, and enable AWS CloudTrail management events.
D.Configure the S3 bucket to use SSE-KMS with the customer-managed key, and enable AWS CloudTrail logging for KMS API calls.
AnswerD

Using SSE-KMS with a customer-managed key ensures data at rest is encrypted with that key, and CloudTrail logs all KMS API calls, providing an audit trail of key usage. This directly meets the encryption and auditing requirements. The Glue job will automatically use the key when reading and writing if permissions are granted.

Why this answer

SSE-KMS with a customer-managed key provides control over the encryption key and enables auditing via CloudTrail, which logs all KMS API calls. This meets both the encryption at rest and auditability requirements. Other options either use AWS-managed keys or do not provide the necessary audit detail.

Exam trap

The trap here is assuming that S3 server access logging or CloudTrail S3 data events capture KMS key usage, when only CloudTrail management events for KMS or KMS key policies provide that visibility.

1039
MCQeasy

A data engineer needs to store JSON documents that are frequently read and written by a web application. The data has a flexible schema and requires low-latency queries on primary key lookups. Which AWS service is MOST suitable?

A.Amazon Redshift
B.Amazon S3
C.Amazon DynamoDB
D.Amazon RDS for MySQL
AnswerC

Amazon DynamoDB stores JSON as native documents and delivers single-digit-millisecond latency on primary key lookups, satisfying the low-latency requirement. Its schemaless key-value design accommodates the flexible schema, while provisioned throughput sustains the frequent reads and writes the web application generates.

Why this answer

Amazon DynamoDB is the most suitable service because it is a NoSQL key-value and document database that provides single-digit millisecond latency for primary key lookups, supports flexible schemas for JSON documents, and is designed for high-throughput read/write workloads from web applications. Its fully managed nature and auto-scaling capabilities align with the requirement for frequent, low-latency queries on a flexible schema.

Exam trap

The trap here is that candidates may confuse Amazon S3's ability to store JSON documents with the need for low-latency primary key lookups, overlooking that S3 is not a database and lacks the indexing and query performance required for frequent, transactional reads and writes.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a columnar data warehouse optimized for complex analytical queries on structured data, not for low-latency primary key lookups on JSON documents with frequent writes. Option B is wrong because Amazon S3 is an object storage service that does not support low-latency primary key lookups or native querying without additional services like Athena or S3 Select, and it is not designed for frequent, transactional read/write operations. Option D is wrong because Amazon RDS for MySQL is a relational database with a fixed schema, requiring schema changes for flexible JSON documents, and while it can handle JSON, it does not match DynamoDB's single-digit millisecond latency for primary key lookups at scale.

1040
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. One of the Glue jobs occasionally fails due to transient network issues. The engineer wants the Step Function to retry the failed Glue job up to 3 times with exponential backoff before failing the entire workflow. Which Step Functions state configuration should be used?

A.Configure a Catch field in the Task state to transition to a Wait state, then loop back to the Glue job Task.
B.Set the Glue job's MaximumRetries parameter to 3 in the job definition.
C.Add a Retry field in the Task state with ErrorEquals matching the Glue job's error, MaxAttempts: 3, and BackoffRate: 2.0.
D.Use a Map state to iterate over the Glue job three times with a Wait state in between.
AnswerC

The Retry field in a Step Functions Task state allows specifying retry behavior for specific errors. Setting MaxAttempts to 3 and BackoffRate to 2.0 implements exponential backoff. This directly meets the requirement to retry up to 3 times with exponential backoff, and is the standard way to handle transient failures in Step Functions.

Why this answer

The Retry field in Step Functions Task states is specifically designed to handle transient errors by retrying the task with configurable attempts and backoff. Setting MaxAttempts to 3 and BackoffRate to 2.0 achieves up to three retries with exponential backoff. This is the standard, declarative approach and avoids custom orchestration logic.

It also allows matching on specific error types, ensuring retries only occur for relevant failures.

Exam trap

The trap here is conflating AWS Glue's internal retry mechanism with Step Functions orchestration, leading to the selection of a Glue-level parameter instead of the Step Functions Retry field.

1041
MCQeasy

A company needs to ingest streaming data from thousands of IoT devices into Amazon S3 for long-term storage and analytics. The data arrives continuously at a rate of 5 MB per second and must be stored in a compressed format to reduce storage costs. The solution should be highly available and require minimal management. Which AWS service should the company use?

A.Amazon Kinesis Data Streams with an AWS Lambda consumer that writes to Amazon S3.
B.Amazon Kinesis Data Firehose with a delivery stream to Amazon S3, enabling compression.
C.Amazon Managed Streaming for Apache Kafka (Amazon MSK) with a Kafka Connect S3 sink connector.
D.AWS Glue streaming ETL job that reads from Amazon Kinesis Data Streams and writes to Amazon S3.
AnswerB

Kinesis Data Firehose is a fully managed service that can ingest streaming data, automatically compress it using GZIP, and deliver it to S3. It requires no code, scales automatically, and provides high availability. It meets the requirements of continuous ingestion, compression, and minimal management, making it the ideal choice.

Why this answer

Amazon Kinesis Data Firehose is designed for streaming data ingestion to destinations like S3. It automatically handles scaling, buffering, compression, and encryption. With compression enabled, data is stored in GZIP format, reducing storage costs.

It requires no infrastructure management, making it the simplest and most cost-effective solution for continuous IoT data ingestion to S3.

Exam trap

The trap here is assuming that Kinesis Data Streams with Lambda is equivalent to Firehose, but Firehose provides built-in compression and delivery to S3 without custom code.

1042
MCQmedium

A data engineer needs to ingest streaming data from thousands of devices sending JSON messages via HTTP POST. The data should be stored in Amazon S3 with minimal latency and also be available for real-time analytics. Which combination of services is MOST appropriate?

A.Amazon DynamoDB with DynamoDB Streams and Lambda.
B.Amazon SQS and AWS Lambda to write to S3.
C.AWS Lambda directly writing to S3 via API Gateway.
D.Amazon API Gateway, Amazon Kinesis Data Streams, and Kinesis Data Firehose.
AnswerD

API Gateway accepts the devices' HTTP POST requests, Kinesis Data Streams provides low-latency ingestion for real-time analytics consumers, and Firehose micro-batches the same stream into S3. This satisfies both the minimal-latency S3 storage and concurrent real-time analytics constraints.

Why this answer

Amazon API Gateway provides a scalable, managed HTTP endpoint that can receive JSON POST requests from thousands of devices and proxy them directly into Amazon Kinesis Data Streams. Kinesis Data Streams ingests the data with low latency and makes it available for real-time analytics via Kinesis Client Library (KCL) applications, AWS Lambda, or Kinesis Data Analytics. Kinesis Data Firehose then delivers the stream to Amazon S3 with minimal buffering (as low as 60 seconds), satisfying both the real-time and storage requirements.

Exam trap

DEA-C01 often tests the misconception that SQS or DynamoDB Streams can replace Kinesis for real-time streaming ingestion, confusing queue-based decoupling with stream-based replayable analytics.

How to eliminate wrong answers

Option A is wrong because DynamoDB Streams captures item-level changes from a DynamoDB table, not raw HTTP POST payloads from devices; it also does not natively deliver to S3 and is not designed for high-volume streaming ingestion. Option B is wrong because SQS is a queue, not a streaming service—it lacks the ordered, replayable, low-latency stream semantics needed for real-time analytics, and Lambda polling SQS adds latency and complexity. Option C is wrong because Lambda writing directly to S3 via API Gateway bypasses a durable, replayable stream, offers no built-in buffering or real-time analytics integration, and can hit Lambda concurrency limits under thousands of concurrent device connections.

1043
MCQmedium

A data engineer needs to store and analyze time-series data from IoT devices. The data volume is 10 GB per day, and the queries are mostly on the most recent 7 days of data. The engineer wants to minimize storage costs while retaining historical data for 1 year. Which combination of AWS services is most cost-effective?

A.Amazon Timestream
B.Amazon DynamoDB with TTL and S3 for archival
C.Amazon Redshift
D.Amazon RDS with MySQL
AnswerA

Timestream is cost-effective for time-series data with automatic storage tiering.

Why this answer

Amazon Timestream is purpose-built for time-series data, offering automatic tiering between in-memory (for recent 7 days) and magnetic stores (for historical data up to 1 year). This matches the query pattern (mostly recent 7 days) and retention requirement (1 year) while minimizing storage costs through its serverless, pay-per-query model. Timestream also supports time-series-specific functions like interpolation and smoothing, making it more efficient than general-purpose databases for this workload.

Exam trap

The trap here is that candidates often choose DynamoDB with TTL and S3 for archival (Option B) because it seems cost-effective, but they overlook the operational complexity and query latency of accessing historical data in S3, which violates the 'minimize storage costs while retaining historical data for 1 year' requirement without considering query patterns.

How to eliminate wrong answers

Option B (DynamoDB with TTL and S3 for archival) is wrong because DynamoDB is optimized for key-value and document workloads, not time-series analytics; TTL only deletes old data, but querying historical data from S3 requires additional services like Athena or Glue, increasing complexity and latency. Option C (Amazon Redshift) is wrong because Redshift is a columnar data warehouse designed for large-scale analytical queries on structured data, but it is over-provisioned and costly for 10 GB/day of time-series data, and its storage and compute are not optimized for time-series-specific operations like downsampling or retention policies. Option D (Amazon RDS with MySQL) is wrong because RDS is a relational database with fixed storage and compute, leading to higher costs for storing 3.65 TB of historical data (10 GB/day × 365 days) and poor query performance on time-series data without built-in time-series features like automatic retention or partitioning.

1044
MCQeasy

A company is using Amazon S3 for data lake storage. They need to query the data directly using SQL without loading it into a database. Which AWS service should be used?

A.Amazon Redshift Spectrum
B.Amazon Athena
C.Amazon EMR
D.AWS Glue
AnswerB

Athena queries data in place in Amazon S3 using standard SQL, requiring no loading into a database. This directly satisfies the requirement to query S3 data lake content with SQL while avoiding ETL or database provisioning.

Why this answer

Amazon Athena is the correct choice because it is a serverless, interactive query service that allows you to analyze data directly in Amazon S3 using standard SQL, without needing to load or transform the data into a database. Athena uses Presto under the hood and supports querying structured, semi-structured, and unstructured data formats (e.g., CSV, JSON, Parquet, ORC) stored in S3, making it ideal for ad-hoc SQL queries on a data lake.

Exam trap

The trap here is that candidates often confuse AWS Glue's data cataloging and ETL capabilities with direct SQL querying, or they assume Redshift Spectrum is a standalone service rather than a feature requiring an existing Redshift cluster, leading them to pick a wrong answer that requires additional infrastructure or is not a query engine.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift Spectrum is a feature of Amazon Redshift that allows querying data in S3 from within a Redshift data warehouse, but it requires an existing Redshift cluster and is not a standalone service for directly querying S3 data without a database. Option C is wrong because Amazon EMR is a big data platform that uses frameworks like Apache Spark, Hive, or Presto for querying S3 data, but it requires provisioning and managing clusters, which adds complexity and is not a serverless SQL-only solution. Option D is wrong because AWS Glue is a serverless data integration service primarily used for ETL (extract, transform, load) jobs and data cataloging, not for directly querying S3 data with SQL; while it can prepare data for Athena, it is not a query engine itself.

1045
Multi-Selecthard

A company uses AWS Glue to transform data stored in S3. The Glue job runs daily and processes data in the range of hundreds of GB. The data engineer wants to optimize the job for cost and performance. Which THREE actions should be taken? (Choose THREE.)

Select 3 answers
A.Store intermediate data in HDFS on Amazon EMR
B.Increase the number of DPUs for the job
C.Reduce the number of DPUs to save cost
D.Use columnar data formats such as Parquet
E.Partition the data by date or other high-cardinality columns
AnswersB, D, E

More DPUs can reduce runtime, improving cost if job runs shorter.

Why this answer

Increasing the number of DPUs (Data Processing Units) for an AWS Glue job can improve performance by enabling parallel processing of large datasets (hundreds of GB). AWS Glue allocates resources in increments of DPUs, where each DPU provides 4 vCPU and 16 GB of memory; scaling out DPUs reduces execution time, which can lower overall cost if the job runs fewer minutes, balancing cost and performance.

Exam trap

The trap here is that candidates mistakenly think reducing DPUs always saves cost, but AWS Glue bills by DPU-hour, so longer runtimes from fewer DPUs can actually increase cost, and the question explicitly asks for both cost and performance optimization.

1046
MCQeasy

A company wants to use Amazon Redshift Spectrum to query data in Amazon S3. The data is in Parquet format and partitioned by date. Which step is required to enable Redshift Spectrum?

A.Load the data into Redshift tables using the COPY command.
B.Create an external schema and external table in the AWS Glue Data Catalog.
C.Create a separate Redshift Spectrum cluster.
D.Copy the data from S3 to Redshift-managed storage.
AnswerB

Redshift Spectrum queries S3 data through external tables registered in the AWS Glue Data Catalog, which you reference via an external schema in the Redshift cluster. This is the mandatory configuration step enabling Spectrum to read the partitioned Parquet data.

Why this answer

Redshift Spectrum allows querying data directly in Amazon S3 without loading it into Redshift. To use Spectrum, you must define an external schema and external table in the AWS Glue Data Catalog (or an external Hive metastore) that points to the S3 location and specifies the Parquet format and partition structure. This enables Redshift to read the data in place using the Spectrum engine.

Exam trap

The trap here is that candidates assume Redshift Spectrum requires a separate cluster or that data must be loaded into Redshift, confusing Spectrum with traditional Redshift ingestion methods like COPY or CTAS.

How to eliminate wrong answers

Option A is wrong because the COPY command loads data into Redshift-managed storage, which bypasses Spectrum's external query capability and incurs storage costs; Spectrum queries data directly from S3 without loading. Option C is wrong because Redshift Spectrum does not require a separate cluster; it runs on the existing Redshift cluster's compute nodes, leveraging the Spectrum layer to access S3. Option D is wrong because copying data from S3 to Redshift-managed storage defeats the purpose of Spectrum, which is to query data in place without moving it.

1047
MCQmedium

A data engineer needs to share an S3 bucket with another AWS account. They want to ensure that the objects in the bucket remain encrypted with SSE-KMS using a customer managed key. What additional step is required for cross-account access?

A.Modify the KMS key policy to grant the target account kms:Decrypt permission
B.Add an IAM policy in the target account to allow kms:Decrypt
C.Disable SSE-KMS encryption on the bucket
D.Add a bucket policy that grants the target account s3:GetObject
AnswerA

Cross-account SSE-KMS access requires both the S3 bucket policy and the KMS key policy to permit the target account. Because the key is customer managed, its key policy must explicitly grant the other account kms:Decrypt, otherwise S3 cannot unwrap the data key for that account's requests.

Why this answer

When using SSE-KMS with a customer managed key, cross-account access requires the KMS key policy to grant the target account's IAM role or user the necessary KMS permissions (kms:Decrypt, and optionally kms:GenerateDataKey). The S3 bucket policy must also grant s3:GetObject, and the target account's IAM policy must allow kms:Decrypt. However, the key policy is the additional step specific to KMS that is not covered by S3 policies alone.

Without it, the target account cannot use the key. Option A is correct because modifying the key policy is essential. Option B is insufficient because the target account's IAM policy cannot override the key policy.

Option C is unnecessary and breaks encryption. Option D provides S3 access but not KMS access.

1048
Multi-Selectmedium

Which TWO options are valid methods to ingest on-premises relational database data into Amazon S3 for analytics? (Choose 2.)

Select 2 answers
A.AWS Snowball Edge
B.AWS Glue ETL job with JDBC connection to source
C.Amazon Kinesis Data Streams with Direct Put
D.AWS Database Migration Service (DMS) with S3 target
E.Amazon AppFlow
AnswersB, D

Glue can read from JDBC and write to S3.

Why this answer

AWS Glue ETL jobs can connect to on-premises relational databases via JDBC, extract data, and write it directly to Amazon S3 in formats like Parquet or ORC. This is a fully managed, serverless approach suitable for batch ingestion and transformation of structured data for analytics.

Exam trap

The trap here is that candidates confuse AWS Glue ETL (which uses JDBC for batch extraction) with Amazon Kinesis (which is for streaming), or they overlook that AWS DMS is a dedicated service for database migration and replication to S3, while Snowball Edge is for offline bulk transfer, not live ingestion.

1049
MCQmedium

A data engineer is using AWS Lake Formation to manage access to a data lake in Amazon S3. The company wants to grant a data analyst read-only access to specific columns in a table stored in the AWS Glue Data Catalog. The analyst should not be able to see other columns or any rows that contain sensitive data. The engineer sets up Lake Formation permissions on the table, granting SELECT on specific columns. However, when the analyst queries the table using Amazon Athena, they can see all columns. What is the most likely reason?

A.The Glue Data Catalog table definition does not include the column-level metadata required for Lake Formation.
B.The analyst has IAM permissions that allow direct access to the S3 bucket, bypassing Lake Formation.
C.Lake Formation column-level permissions are not supported for Athena queries.
D.The analyst's IAM role lacks the lakeformation:GetDataAccess permission.
AnswerB

Lake Formation uses a permission model that requires the principal to have no direct IAM access to the underlying S3 data. If the analyst has IAM permissions to read the S3 bucket, they can bypass Lake Formation's column-level restrictions and access all data directly. To enforce Lake Formation permissions, the analyst's IAM policy must not allow direct S3 access, and all access should go through Lake Formation-enabled services.

Why this answer

Lake Formation enforces fine-grained access control only when users do not have direct IAM permissions to the underlying data. If the analyst has IAM permissions to read the S3 bucket, they can bypass Lake Formation and access all columns and rows. To enforce column-level security, the analyst's IAM policy must not allow direct S3 access; all data access must be mediated through Lake Formation.

The other options are either incorrect or would produce different errors.

Exam trap

The trap here is assuming that Lake Formation permissions alone are sufficient, while ignoring that direct IAM access to S3 can bypass those permissions entirely.

1050
MCQeasy

A company needs to store files that are accessed by multiple EC2 instances in a VPC. The files must be concurrently accessible and durable. Which storage solution should the data engineer choose?

A.Amazon EC2 instance store
B.Amazon Simple Storage Service (Amazon S3)
C.Amazon Elastic Block Store (Amazon EBS)
D.Amazon Elastic File System (Amazon EFS)
AnswerD

Amazon EFS provides a shared, elastic NFS file system that many EC2 instances mount concurrently across Availability Zones, delivering the concurrent access and durability the scenario demands. Instance store and EBS volumes attach to a single instance, so they cannot satisfy the multi-instance requirement.

Why this answer

Amazon EFS provides a fully managed, scalable, and elastic NFS file system that can be concurrently accessed by multiple EC2 instances across multiple Availability Zones. It is designed for high durability (11 nines of durability) and automatically replicates data across multiple AZs within a region, meeting the requirements for concurrent access and durability.

Exam trap

The trap here is that candidates often confuse Amazon EBS Multi-Attach with a general-purpose shared file system, but EBS Multi-Attach is limited to specific io1/io2 volumes, requires cluster-aware applications, and does not provide the POSIX file system semantics or cross-AZ durability that EFS offers.

How to eliminate wrong answers

Option A is wrong because EC2 instance store provides ephemeral block storage that is physically attached to the host; it is not durable (data is lost on instance stop/termination) and cannot be shared concurrently across multiple EC2 instances. Option B is wrong because Amazon S3 is an object storage service, not a file system; it does not support standard file-level locking or NFS/SMB protocols required for concurrent file access from multiple EC2 instances without additional gateways or software. Option C is wrong because Amazon EBS provides block-level storage volumes that can only be attached to a single EC2 instance at a time (except for multi-attach EBS io1/io2 volumes, which are limited to specific instance types and have strict constraints, not a general solution for concurrent file access).

Page 13

Page 14 of 18

Page 15