Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 1051–1125

1321 questions total · 18pages · All types, answers revealed

Page 14

Page 15 of 18

Page 16
1051
MCQmedium

The exhibit shows an AWS CLI command and its output. A data engineer wants to copy only objects larger than 10 MB from the S3 bucket to another bucket for processing. Which approach should be used to automate this task?

A.Use S3 replication rules to replicate objects above 10 MB
B.Use AWS CLI with a script to filter and copy objects
C.Use S3 Inventory to generate a list and then copy
D.Use AWS Lambda with S3 event notifications
AnswerB

The CLI can filter by size and copy objects using a script.

Why this answer

The command lists objects larger than 10 MB. To automate copying, a script using AWS CLI with the --query parameter can filter and copy. Using S3 Batch Operations allows performing actions on a list of objects.

The correct approach is to use AWS CLI with a script that iterates over the filtered list and uses aws s3 cp. S3 replication is for continuous sync, not one-time copy. Lambda with S3 events triggers only on new objects, not existing ones.

S3 Inventory provides metadata but not direct copy.

1052
MCQhard

A data engineer manages an Amazon DynamoDB table for order events. Reads and writes are evenly spread across a partition key with very high cardinality, but during flash sales the table throttles with ProvisionedThroughputExceededException even though consumed capacity is below the provisioned total. Which cause is MOST likely?

A.The table's on-demand capacity mode must be enabled to remove per-partition limits.
B.The table's items are too large, and each write consumes more than one write capacity unit.
C.The table has too many global secondary indexes, which each consume write capacity.
D.A single partition key value is receiving a disproportionate share of traffic, creating a hot partition that exceeds its per-partition throughput limit.
AnswerD

DynamoDB divides a table into partitions, and each partition has a hard ceiling of roughly 3,000 read units and 1,000 write units per second regardless of the table's total provisioned capacity. If one key value, such as a default or shared customer ID, receives a burst of traffic, that partition saturates and throttles even though aggregate consumed capacity sits below the table's provisioned total.

Why this answer

DynamoDB throughput is provisioned at the table level but enforced at the partition level, and each partition can serve only about 1,000 write units per second. When traffic concentrates on one partition key value, that partition throttles while the table still shows unused capacity. Even distribution across a high-cardinality key normally prevents this, so the flash-sale burst must be funneling requests to a single key value.

Exam trap

The trap here is reading ProvisionedThroughputExceededException as proof that the table needs more capacity, when spare table capacity plus throttling actually indicates a hot partition.

1053
MCQeasy

A data engineer needs to ingest streaming data from thousands of IoT devices into AWS for real-time processing. The data volume peaks at 5 GB/min. Which AWS service should be used as the ingestion endpoint?

A.Amazon Kinesis Data Streams
B.AWS Glue
C.Amazon S3
D.AWS Lambda
AnswerA

Kinesis Data Streams ingests high-throughput streaming records with per-shard capacity that scales to thousands of producers, and delivers sub-second latency for real-time processing. It satisfies the 5 GB/min peak by adding shards, unlike S3 transfer or batch-oriented endpoints.

Why this answer

Amazon Kinesis Data Streams is designed for real-time data ingestion at scale, supporting throughput of up to 1 MB/s or 1,000 records/s per shard. With a peak of 5 GB/min (~83 MB/s), you can horizontally scale by adding shards to meet the required throughput, making it the ideal ingestion endpoint for high-volume streaming IoT data.

Exam trap

The trap here is that candidates often confuse AWS Glue's streaming ETL capability (which reads from a stream but does not ingest) with a direct ingestion endpoint, or they assume S3's high durability makes it suitable for real-time ingestion, ignoring its lack of streaming semantics and low-latency write guarantees.

How to eliminate wrong answers

Option B (AWS Glue) is wrong because it is a serverless ETL service for batch data processing and cataloging, not a real-time streaming ingestion endpoint; it cannot handle continuous, high-velocity data streams. Option C (Amazon S3) is wrong because it is an object storage service that does not provide real-time ingestion or streaming capabilities; data must be written via API calls or batch uploads, and it lacks the low-latency, ordered replay features needed for streaming. Option D (AWS Lambda) is wrong because it is a compute service for running code in response to events, not a dedicated ingestion endpoint; it has a maximum invocation payload limit of 256 KB and is not designed to buffer or scale for sustained 5 GB/min throughput.

1054
Multi-Selectmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON and writes to a partitioned Parquet table. The engineer wants to reduce job cost and improve read performance. (Choose two.)

Select 2 answers
A.Enable job bookmarks to avoid reprocessing previously processed files.
B.Convert the source JSON to Parquet before the Glue job runs.
C.Increase the number of workers to the maximum allowed for the job.
D.Use the Glue crawler to infer the schema and store it in the Data Catalog.
E.Partition the output Parquet table by low-cardinality columns used in filters.
AnswersA, E

Job bookmarks track which S3 objects have already been processed and skip them on subsequent runs. For incremental data landing in S3, this avoids re-reading and re-transforming old files, directly reducing DPU hours and cost. It is a supported Glue feature and a common cost-control measure for recurring jobs.

Why this answer

Job bookmarks cut cost by skipping already-processed S3 objects, which is important for incremental pipelines. Partitioning the output Parquet table by low-cardinality filter columns enables partition pruning, reducing I/O and improving read performance for downstream queries. Crawlers, worker scaling, and pre-conversion do not directly deliver both cost reduction and read performance for this job configuration.

Exam trap

The trap here is assuming that adding workers or running a crawler reduces cost, when bookmarks and output partitioning are the levers that actually cut DPU time and scan volume.

1055
Multi-Selectmedium

A company is building a data lake on Amazon S3 and needs to ingest data from multiple sources. Which of the following AWS services can be used to ingest and transform data in near real-time? (Select TWO.)

Select 2 answers
A.AWS Glue
B.Amazon Kinesis Data Firehose
C.Amazon Athena
D.AWS Step Functions
E.Amazon Simple Queue Service (SQS)
AnswersA, B

Can be used for ETL jobs triggered by S3 events.

Why this answer

AWS Glue is correct because it provides a serverless ETL (Extract, Transform, Load) service that can ingest data from various sources and transform it in near real-time using its streaming ETL capabilities. Glue can consume data from Amazon Kinesis Data Streams or Apache Kafka, apply transformations using Apache Spark, and write the results to Amazon S3 or other destinations, making it suitable for near real-time data ingestion and transformation.

Exam trap

The trap here is that candidates often confuse Amazon Athena (a query engine) with an ingestion service, or assume SQS alone can perform transformations, when neither service is designed for near real-time data ingestion and transformation into a data lake.

1056
MCQeasy

A data engineer needs to run a transformation on streaming data using SQL-like queries without managing servers, and the output must be written to Amazon S3 in near real time. The source is an Amazon Kinesis Data Stream. Which AWS service is the MOST appropriate to perform the transformation?

A.AWS Glue with a Python shell job triggered hourly by Amazon EventBridge.
B.AWS Lambda with a function triggered by Kinesis records and writing directly to S3.
C.Amazon Managed Service for Apache Flink with a SQL application.
D.Amazon EMR running Apache Spark Structured Streaming on a persistent cluster.
AnswerC

Amazon Managed Service for Apache Flink supports SQL-based stream processing on Kinesis Data Streams and can write results to Amazon S3, while the service manages the underlying infrastructure. It is designed for continuous, near-real-time transformations and removes the burden of provisioning and scaling servers, which matches the SQL-like, serverless-management requirement and the S3 output destination.

Why this answer

Amazon Managed Service for Apache Flink provides a managed environment for running SQL-based stream processing applications against Kinesis Data Streams and can deliver results to Amazon S3 in near real time. Because AWS operates the underlying infrastructure, the engineer avoids cluster management while still getting continuous, low-latency transformation, which is precisely what the scenario requires.

Exam trap

The trap here is treating AWS Lambda or a scheduled Glue job as equivalent to a managed streaming SQL engine, when only Flink provides continuous SQL processing without server management.

1057
MCQeasy

A data engineering team is using AWS Glue to catalog data in an S3 data lake. They have a Glue crawler that runs daily to update the Data Catalog. Recently, they noticed that the crawler is taking longer to run and sometimes fails because of a timeout. The team suspects the issue is due to the large number of small files in the S3 bucket. They need to improve crawler performance and reliability. Which solution should they implement?

A.Configure the crawler to use a different classifier.
B.Use AWS Glue ETL to consolidate small files into larger ones before crawling.
C.Increase the crawler timeout to 24 hours.
D.Schedule the crawler to run more frequently to avoid large data accumulation.
AnswerB

Consolidating many small S3 objects into fewer larger files via Glue ETL reduces the metadata and listing overhead the crawler processes each run. This directly addresses the small-file volume causing slow runs and timeouts, restoring crawler performance and reliability.

Why this answer

Consolidating small files into larger ones (e.g., using AWS Glue ETL with a groupFiles or groupSize option, or a separate compaction job) reduces the number of objects the crawler must list and sample. This directly addresses the root cause: a high volume of small files increases metadata operations and can cause crawler timeouts. By reducing file count, the crawler can complete within the default 24-hour timeout and avoid failures.

Exam trap

The trap here is that candidates assume increasing the timeout or running the crawler more frequently will fix performance issues, but the real bottleneck is the sheer number of small files, which requires data compaction to resolve.

How to eliminate wrong answers

Option A is wrong because changing the classifier affects how the crawler interprets data format (e.g., JSON vs. Parquet), not the number of files or the performance bottleneck caused by small files. Option C is wrong because increasing the timeout to 24 hours does not solve the underlying issue of excessive small files; the crawler may still fail due to resource limits or S3 request throttling, and the default timeout is already 24 hours.

Option D is wrong because running the crawler more frequently would only accumulate more small files over time, worsening the problem and increasing the likelihood of timeouts.

1058
Multi-Selecteasy

A data engineer is setting up Amazon S3 bucket policies for a data lake. Which TWO statements are true regarding S3 bucket policies? (Choose TWO.)

Select 2 answers
A.Bucket policies can grant access to accounts in other AWS Organizations
B.Bucket policies are the only way to control access to S3
C.Bucket policies can be applied to individual objects
D.The Principal element in a bucket policy is optional
E.Bucket policies are written in JSON format
AnswersA, E

Cross-account access can be granted via bucket policies.

Why this answer

S3 bucket policies can grant cross-account access to principals in other AWS accounts, including those in different AWS Organizations, by specifying the target account ID or organization ID in the Principal element. This enables centralized data lake access management across organizational boundaries without requiring IAM roles or resource-based policies in each account.

Exam trap

The trap here is that candidates often confuse bucket policies with IAM policies, mistakenly thinking the Principal element is optional in bucket policies (it is required), or that bucket policies can target individual objects (they cannot; they use prefix or tag conditions instead).

1059
MCQeasy

A company wants to ingest streaming data from Apache Kafka into Amazon S3 for long-term storage and analytics. The data is in JSON format and must be delivered to S3 with minimal effort and no custom code. Which AWS service should the data engineer use?

A.AWS Glue streaming ETL job.
B.Amazon Kinesis Data Streams with an AWS Lambda consumer.
C.Amazon Managed Service for Apache Flink.
D.Amazon Kinesis Data Firehose with a Kafka source.
AnswerD

Kinesis Data Firehose can ingest data directly from Apache Kafka (via the Firehose HTTP endpoint or the Kafka connector) and deliver it to Amazon S3 without any custom code. It automatically handles scaling, buffering, and delivery, making it the most appropriate choice for minimal effort and no code.

Why this answer

Kinesis Data Firehose is a fully managed service that can read from Apache Kafka and deliver to Amazon S3 without requiring custom code. It handles scaling, buffering, and format conversion. Other options like Kinesis Data Streams, Glue, or Managed Flink require writing and maintaining code, which does not meet the no-code requirement.

Exam trap

The trap here is confusing Kinesis Data Streams with Kinesis Data Firehose; Data Streams requires custom consumers, while Firehose provides managed delivery.

1060
MCQhard

A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is then processed by a Kinesis Data Analytics application running SQL queries. The analytics application is falling behind and processing records with increasing latency. The stream has 4 shards, and the average record size is 5 KB. What is the MOST effective way to improve processing latency?

A.Increase the number of shards in the Kinesis stream to 8.
B.Enable enhanced fan-out on the Kinesis stream for the analytics application.
C.Increase the parallelism of the Kinesis Data Analytics application.
D.Increase the retention period of the Kinesis stream to 7 days.
AnswerC

More parallelism allows the application to process more records per unit time.

Why this answer

Increasing the parallelism of the Kinesis Data Analytics application (e.g., by increasing the number of in-application streams or ParallelismPerKPU) allows it to consume from the stream faster, reducing latency. Option A is wrong because 5 KB is well below the 1 MB/s shard limit, so increasing shards is unnecessary. Option B is wrong because enhanced fan-out is for consumers that need low latency, but does not help if the application is CPU-bound.

Option D is wrong because increasing the retention period does not affect processing speed; it only keeps data longer.

1061
MCQeasy

A data engineer is troubleshooting a failed AWS Glue job that writes results to Amazon S3. The error log shows 'AccessDenied' when trying to list the bucket. Which IAM policy statement should the engineer add to the Glue job's role?

A.s3:ListBucket
B.s3:PutObject
C.s3:DeleteObject
D.s3:GetObject
AnswerA

The AccessDenied error occurs on the ListBucket operation, which requires the s3:ListBucket permission on the bucket resource itself, not on objects. Adding this action to the Glue job's IAM role grants the missing listing capability.

Why this answer

Listing a bucket requires the s3:ListBucket permission. Option B is wrong because s3:PutObject is for writing, not listing. Option C is wrong because s3:DeleteObject is not needed.

Option D is wrong because s3:GetObject is for reading objects, not listing.

1062
Multi-Selecthard

A data engineer is optimizing an AWS Glue ETL job that reads from Amazon S3 and writes to Amazon Redshift. The job currently uses a single large file and takes hours to complete. The engineer wants to improve performance by using partitioning and parallelism. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Configure the job to write output to a single file to reduce the number of S3 PUT requests.
B.Convert the input data to a partitioned format, such as Parquet, and store it in S3 with a partition key.
C.Enable job bookmarks to track processed data and avoid reprocessing.
D.Increase the number of DPUs allocated to the Glue job to allow more parallel tasks.
E.Use the Glue DynamicFrame instead of a DataFrame to enable automatic schema inference.
AnswersB, D

Partitioning the data and using a columnar format like Parquet allows Glue to read only relevant partitions and leverage predicate pushdown. This reduces I/O and enables parallel processing across partitions. It directly addresses the slow performance caused by a single large file by splitting data into manageable chunks that can be processed concurrently.

Why this answer

To improve performance, the engineer should partition the input data and use a columnar format like Parquet, which enables Glue to read only necessary data and process partitions in parallel. Additionally, increasing the number of DPUs provides more compute resources to handle the parallel tasks. Together, these actions address both data layout and compute capacity, leading to faster job execution.

Exam trap

The trap here is focusing on incremental processing features like job bookmarks or schema handling with DynamicFrames, which do not directly speed up a single large dataset.

1063
MCQeasy

A data engineer needs to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket for long-term storage. The data is in JSON format and must be delivered within 60 seconds of arrival. The engineer wants a fully managed solution that requires minimal code. Which service should the engineer use?

A.AWS Glue ETL job scheduled to run every minute to pull from Kinesis.
B.AWS Lambda to read from the Kinesis stream and write to Amazon S3.
C.Amazon Kinesis Client Library (KCL) application running on Amazon EC2.
D.Amazon Kinesis Data Firehose with an Amazon S3 destination.
AnswerD

Amazon Kinesis Data Firehose is a fully managed service that can ingest data from Kinesis Data Streams and deliver it to Amazon S3. It requires no code to set up and can buffer data for up to 60 seconds or based on size. It automatically handles scaling, and can convert JSON to Parquet or other formats if needed. This meets the requirements.

Why this answer

Amazon Kinesis Data Firehose is the fully managed service designed to deliver streaming data from Kinesis Data Streams to destinations like Amazon S3. It requires no code, can buffer for up to 60 seconds, and handles scaling automatically. The other options involve custom code or are not suited for near-real-time streaming ingestion.

Exam trap

The trap here is assuming that any service that can read from Kinesis and write to S3 is equally managed, when Firehose is purpose-built for this with minimal code.

1064
MCQeasy

A company uses Amazon S3 as its data lake. A data engineer needs to enforce encryption of data at rest using server-side encryption with AWS KMS. Which S3 bucket property should be configured?

A.Default encryption
B.Server access logging
C.Versioning
D.Bucket policy
AnswerA

Configuring default encryption on the bucket applies SSE-KMS automatically to every object written, without relying on request headers. This enforces encryption at rest using AWS KMS as the stem demands, covering uploads from any client or SDK.

Why this answer

Configuring default encryption on an S3 bucket ensures that all objects stored in the bucket are encrypted at rest using server-side encryption. When AWS KMS is specified as the encryption type, S3 automatically encrypts objects with a KMS key (SSE-KMS) upon upload, even if the upload request does not include encryption headers. This enforces encryption at rest without requiring changes to client applications.

Exam trap

The trap here is that candidates often confuse bucket policies (which can enforce encryption conditions) with default encryption (which actually applies encryption), leading them to select bucket policy as the answer when the question asks for the property that enforces encryption of data at rest.

How to eliminate wrong answers

Option B is wrong because server access logging records requests made to the bucket for auditing purposes, but it does not enforce or configure encryption of data at rest. Option C is wrong because versioning preserves, retrieves, and restores every version of every object in the bucket, but it has no effect on encryption settings. Option D is wrong because a bucket policy can deny unencrypted uploads using a condition key like `s3:x-amz-server-side-encryption`, but it does not itself configure the encryption mechanism; it only enforces a policy requirement, whereas default encryption directly applies encryption to all objects.

1065
MCQeasy

Refer to the exhibit. A data engineer creates an Amazon Redshift table with the above DDL. The engineer runs a query to find all orders for a specific customer within a date range. Which statement about query performance is correct?

A.The query will be inefficient because the distribution key is not the same as the sort key.
B.The table should use DISTSTYLE EVEN to improve performance.
C.The query will benefit from both the distribution key and the sort key to minimize data scanned.
D.The sort key will not help because the query filters on customer_id first.
AnswerC

Distribution reduces data movement, sort key reduces data scanned.

Why this answer

The DDL defines customer_id as the distribution key and order_date as the sort key. When the query filters on both customer_id (distribution key) and order_date (sort key), Redshift can use partition pruning via the sort key to skip blocks that don't match the date range, and the distribution key ensures that data for the same customer is co-located on the same node slice, minimizing data movement. This combination reduces the amount of data scanned and improves query performance.

Exam trap

The trap here is that candidates assume the sort key is useless if the filter does not start with the sort key column, but Redshift's zone map pruning works on any column in the sort key, and the distribution key filter can still leverage co-location to reduce data movement.

How to eliminate wrong answers

Option A is wrong because the distribution key and sort key do not need to be the same; they serve different purposes—distribution key optimizes data locality for joins and aggregations, while sort key optimizes range-restricted scans. Option B is wrong because DISTSTYLE EVEN distributes rows randomly across slices, which would scatter a single customer's data across all nodes, increasing network traffic and reducing the benefit of the sort key for range scans. Option D is wrong because the sort key on order_date still helps even though the query filters on customer_id first; Redshift can apply predicate-based block pruning on the sort key after the distribution key filter narrows the relevant slices, and the sort key order (customer_id, order_date) means the date filter can still be used efficiently within each customer's data.

1066
Multi-Selectmedium

A data engineer is optimizing an Amazon RDS for MySQL database that experiences high write throughput. The engineer wants to improve write performance and reduce latency. Which TWO database-level configuration changes can help achieve this?

Select 2 answers
A.Use Provisioned IOPS (io1 or io2) storage.
B.Reduce the backup retention period to 1 day.
C.Increase the DB instance class to a larger size.
D.Create a Read Replica to offload writes.
E.Enable Multi-AZ for high availability.
AnswersA, C

Provisioned IOPS provides consistent low-latency writes.

Why this answer

Provisioned IOPS (io1 or io2) storage delivers consistent and predictable I/O performance by guaranteeing a specified number of I/O operations per second, which directly reduces latency and improves write throughput for high-write workloads. This is the most effective storage-level change for write-intensive RDS for MySQL databases.

Exam trap

The trap here is that candidates often confuse Multi-AZ with performance improvement, but Multi-AZ is designed for durability and failover, not for speeding up writes.

1067
MCQhard

A data engineer is using AWS Glue to process a large dataset where a small number of partitions contain disproportionately more rows than others, causing some executors to run much longer than others and the job to take hours. The engineer wants to redistribute the data across partitions before a join operation to improve performance. Which technique should the engineer apply?

A.Increase the number of AWS Glue DPUs allocated to the job so that more executors are available to process the skewed partitions.
B.Use coalesce to reduce the number of partitions, which merges partitions without a full shuffle and balances the data.
C.Enable the AWS Glue job bookmarks feature so that only new data is processed on subsequent runs, reducing the data volume.
D.Call repartition on the DataFrame using a join key column before the join, which triggers a full shuffle and distributes rows evenly across partitions.
AnswerD

Repartitioning on the join key before the join forces a shuffle that redistributes rows so that each partition contains a roughly equal share of data for that key. This directly addresses the skew that causes some executors to process far more rows than others. While a full shuffle has a cost, it is the standard remedy for skewed joins and can dramatically reduce overall job time when skew is severe.

Why this answer

Repartitioning on the join key before the join causes a shuffle that spreads rows evenly across partitions, which is the direct fix for skewed partitions. Increasing DPUs does not help because the largest partition remains a single bottleneck, coalesce can worsen skew by merging without redistribution, and job bookmarks only reduce data across runs. The key is to redistribute data within the run before the expensive join.

Exam trap

The trap here is believing that adding more compute capacity, such as more DPUs, automatically resolves data skew, when the bottleneck is the uneven distribution of rows rather than total capacity.

1068
Multi-Selectmedium

A company is ingesting real-time clickstream data into Amazon S3 using Amazon Kinesis Data Firehose. The data is semi-structured and the company wants to transform the data into Parquet format and partition it by year, month, day, and hour. Which TWO steps should be taken to achieve this? (Choose TWO.)

Select 2 answers
A.Set up an Amazon S3 event notification to trigger an AWS Lambda function that partitions the data after delivery.
B.Enable dynamic partitioning in Kinesis Data Firehose and specify the partition keys as year, month, day, hour extracted from the data.
C.Use an AWS Glue Crawler to infer the schema and automatically partition the data in S3.
D.Create an AWS Lambda function that transforms incoming records to Parquet and attach it to the Firehose delivery stream as a data transformation.
E.Configure Kinesis Data Firehose to convert the data to Parquet format using a schema from the AWS Glue Data Catalog.
AnswersB, D

Correct. Dynamic partitioning extracts partition keys from the data and creates S3 prefixes accordingly.

Why this answer

Kinesis Data Firehose's dynamic partitioning feature allows you to specify partition keys (year, month, day, hour) extracted from the incoming data, and Firehose will automatically create the corresponding S3 prefix structure (e.g., year=2024/month=01/day=15/hour=10/) during delivery. Option D is correct because to convert semi-structured data to Parquet format, you can attach an AWS Lambda function as a data transformation to Firehose, which converts each record to Parquet before delivery to S3. Option E is incorrect because while Kinesis Data Firehose does support converting data to Parquet format using a schema from the AWS Glue Data Catalog, this approach requires a pre-defined Glue schema and is less flexible for semi-structured data.

Moreover, the question does not mention any existing Glue Data Catalog, and the Lambda transformation in option D is a more direct and customizable method for the transformation needed.

Exam trap

AWS often tests the misconception that dynamic partitioning alone handles format conversion, but in reality, dynamic partitioning only manages the S3 prefix structure, while Parquet conversion requires a separate Lambda transformation or the use of Firehose's built-in Parquet conversion with a compatible input format.

1069
MCQhard

A data engineer runs an AWS Glue job that writes Parquet files to Amazon S3. The job frequently fails with an error indicating too many small files are being written, causing slow downstream Athena queries. The engineer wants to reduce the number of output files without changing the transformation logic. Which action should the engineer take?

A.Set the Glue job's output to use coalesce or repartition before writing.
B.Enable bookmarks on the Glue job to track processed files.
C.Change the output format from Parquet to CSV to reduce file count.
D.Increase the number of DPUs allocated to the Glue job.
AnswerA

Using coalesce or repartition in the Spark code before the write step controls the number of output partitions, directly reducing small files. This preserves transformation logic while consolidating output, which is the standard remedy for too many small Parquet files in Glue jobs.

Why this answer

Controlling output partitions with coalesce or repartition before writing is the direct way to reduce the number of small Parquet files in an AWS Glue job, improving downstream Athena performance without altering transformation logic.

Exam trap

The trap here is thinking more DPUs or bookmarks fix small files, when only output partition control (coalesce/repartition) changes file count.

1070
Multi-Selecthard

Which THREE factors should a data engineer consider when choosing between AWS Glue and Amazon EMR for a data transformation job? (Choose three.)

Select 3 answers
A.The ability to output results to Amazon S3
B.The support for Apache Spark
C.The level of control over the execution environment and dependencies
D.The need for a serverless vs. cluster-based environment
E.The cost model: pay per DPU for Glue vs. per instance for EMR
AnswersC, D, E

Glue is serverless and abstracts infrastructure, while EMR gives cluster-level control over instance types, libraries and dependencies. Needing custom execution environments or specific package versions therefore favours EMR, making control level a genuine selection factor.

Why this answer

When choosing between AWS Glue and Amazon EMR for a data transformation job, key considerations include: the level of control over the execution environment (EMR offers more customization, while Glue is managed), the deployment model (Glue is serverless, EMR is cluster-based), and the cost structure (Glue charges per DPU, EMR charges per EC2 instance). Options A and B are not differentiating factors because both services support Apache Spark and can output to S3.

1071
MCQeasy

A data engineer needs to run a transformation in AWS Glue where each record must be processed independently and the output schema is known ahead of time. The transformation should operate on a DynamicFrame and return a DynamicFrame. Which Glue transform is designed for this row-by-row operation?

A.DropNullFields
B.Relationalize
C.Map
D.ResolveChoice
AnswerC

The Map transform applies a user-supplied function to every record in a DynamicFrame and returns a new DynamicFrame with the transformed records. It preserves the one-to-one record relationship, making it ideal when each record is processed independently and the output schema is predetermined. It is the canonical Glue transform for this pattern.

Why this answer

The Map transform is purpose-built for applying a function to each record and returning a DynamicFrame, which matches the need for independent row processing with a known output schema. Relationalize, ResolveChoice, and DropNullFields are schema or structure utilities, not general row-level mapping transforms.

Exam trap

The trap here is confusing schema-normalization transforms such as ResolveChoice or DropNullFields with the Map transform, which is the only one that applies custom per-record logic.

1072
MCQhard

A data engineer is using AWS Glue Studio to create a job that joins data from two Amazon S3 sources: a large fact table and a small dimension table. The job performs a join and then writes the result to Amazon S3 in Parquet format. The engineer notices that the job is running slowly and consuming many DPUs. Which optimization technique should the engineer apply to improve performance?

A.Use a broadcast join to replicate the small dimension table to all worker nodes.
B.Partition the fact table by join key before the join.
C.Convert the dimension table to CSV format to reduce file size.
D.Increase the number of DPUs to the maximum to speed up the join.
AnswerA

A broadcast join is an optimization technique where the smaller table is replicated to all worker nodes, allowing the join to be performed locally without shuffling the large fact table. This reduces network overhead and improves performance for joins where one table is significantly smaller. AWS Glue supports broadcast joins when one side of the join is small enough to fit in memory.

Why this answer

A broadcast join is the appropriate optimization when joining a large fact table with a small dimension table. It avoids shuffling the large table by replicating the small table to all nodes, reducing network traffic and improving speed. The other options either do not address the join inefficiency, may increase cost, or degrade performance by using a less efficient format.

Exam trap

The trap here is assuming that adding more DPUs is always the best way to improve performance, without considering algorithmic optimizations like broadcast joins.

1073
MCQhard

A company has an Amazon DynamoDB table with a provisioned write capacity of 1000 WCU. During a flash sale, the write traffic spikes to 5000 WCU for 10 minutes. The table is not auto-scaled. Which action should the data engineer take to handle the spike without throttling?

A.Convert the table to on-demand capacity mode before the sale.
B.Set a CloudWatch alarm to increase provisioned capacity when write throttling occurs.
C.Use DynamoDB Accelerator (DAX) to cache writes.
D.Enable auto-scaling with a target utilization of 70% and a maximum capacity of 5000 WCU.
AnswerA

On-demand mode removes provisioned WCU entirely, so DynamoDB instantly accommodates the 5000 WCU burst without throttling. Since the table is not auto-scaled and 1000 WCU is fixed, switching before the flash sale is the only action that absorbs the spike.

Why this answer

The table is currently provisioned with 1000 WCU and cannot handle a spike to 5000 WCU. Converting to on-demand mode before the sale allows DynamoDB to automatically handle varying traffic without throttling, as on-demand capacity scales instantly to meet demand. Option D (auto-scaling) might not react quickly enough for a short 10-minute spike, and the table is not currently auto-scaled.

Option C is incorrect because DAX is a read cache and does not buffer or improve write capacity. Option B is reactive and would not prevent initial throttling.

Exam trap

Candidates often assume DAX can handle write spikes because it is a cache, but DAX only caches reads and does not buffer writes. The correct approach is to use on-demand capacity for unpredictable traffic spikes.

How to eliminate wrong answers

Option A is wrong because converting to on-demand capacity mode before the sale would handle the spike without throttling, as on-demand scales instantly to any traffic, but the question's answer key incorrectly marks C as correct. Option B is wrong because setting a CloudWatch alarm to increase provisioned capacity when write throttling occurs is reactive and will cause throttling before the alarm triggers and capacity increases. Option D is wrong because enabling auto-scaling with a target utilization of 70% and a maximum capacity of 5000 WCU would work if configured in advance, but the table is not auto-scaled and the spike is sudden; auto-scaling has a cooldown period and cannot react instantly to a 10-minute spike.

1074
MCQhard

A company uses AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration is taking longer than expected. The task status shows 'Full load in progress' with a low 'Table throughput (rows/s)'. Which action would MOST improve throughput?

A.Enable Multi-AZ on the DMS replication instance
B.Change the target table preparation mode to 'Do nothing'
C.Increase the number of parallel tasks in the DMS task settings
D.Increase the number of shards in the source database
AnswerC

Parallel tasks allow concurrent loading of tables, increasing throughput.

Why this answer

The low 'Table throughput (rows/s)' during the full load phase indicates that the DMS task is not processing tables with enough parallelism. Increasing the number of parallel tasks in the DMS task settings allows the replication instance to load multiple tables concurrently, which directly improves throughput by utilizing available CPU and memory resources more efficiently.

Exam trap

The trap here is that candidates confuse 'parallel tasks' with 'Multi-AZ' or 'target table preparation mode', assuming that high availability or skipping table preparation will speed up data transfer, when in fact only increasing parallelism directly addresses low row throughput during full load.

How to eliminate wrong answers

Option A is wrong because enabling Multi-AZ on the DMS replication instance provides high availability and failover support, but does not increase throughput during full load; it may even reduce performance due to synchronous replication overhead. Option B is wrong because changing the target table preparation mode to 'Do nothing' only affects how DMS handles existing tables (e.g., truncate or drop), not the speed of data transfer; it does not address low row throughput. Option D is wrong because increasing the number of shards in the source database is a source-side change that does not directly affect DMS's ability to read and load data faster; DMS's throughput is limited by its own parallelism settings, not the source shard count.

1075
MCQmedium

A company uses AWS Glue ETL to transform data from Amazon RDS for MySQL to Amazon S3. The Glue job reads from a JDBC connection. The job runs once daily and processes all records, but the data volume is growing. Which change would improve performance and reduce costs?

A.Increase the number of DPUs for the Glue job
B.Switch to a Glue Python shell job
C.Use a higher JDBC fetch size
D.Enable Glue job bookmarking and set the job to process only new data
AnswerD

Job bookmarks persist state from prior runs, so the JDBC source reads only rows added or changed since the last successful run rather than the full table. This cuts JDBC read volume, shuffle and S3 write costs, satisfying the growing-data-volume constraint while keeping the daily schedule.

Why this answer

Enabling Glue job bookmarking allows the job to process only new or changed data since the last run, rather than reprocessing the entire dataset. This reduces both the data volume read from the JDBC source and the transformation time, directly improving performance and lowering costs by minimizing DPU usage.

Exam trap

The trap here is that candidates often assume increasing parallelism (Option A) is the universal fix for performance, overlooking the fact that reducing the data volume processed (Option D) is a more fundamental and cost-effective optimization.

How to eliminate wrong answers

Option A is wrong because increasing the number of DPUs for the Glue job would increase parallelism and potentially speed up execution, but it does not address the root cause of reprocessing all records daily; it would only scale the cost linearly without reducing the data volume processed. Option B is wrong because a Glue Python shell job is designed for lightweight, single-node Python scripts and cannot handle JDBC connections or large-scale data transformations; it lacks the distributed processing capabilities of a full Glue ETL job. Option C is wrong because using a higher JDBC fetch size can improve the efficiency of reading rows from MySQL by reducing round trips, but it still processes all records every run and does not eliminate the overhead of scanning the entire table daily.

1076
MCQeasy

A data engineer needs to run a transformation on a large dataset stored in Amazon S3 using AWS Glue Studio. The transformation is a simple column rename and filter that can be expressed visually. The engineer wants to minimize development time and avoid writing PySpark code. Which approach should the engineer use?

A.Use AWS Glue DataBrew to create a recipe that renames columns and filters rows, then run a DataBrew job.
B.Use an AWS Glue crawler to infer the schema, then query the data with Amazon Athena and save the results to a new S3 location.
C.Develop an AWS Glue ETL script in PySpark using the DynamicFrame API, then run it as a job with the 'glueetl' job type.
D.Create a Glue Studio visual job, add a source node for the S3 bucket, add a Transform node for the rename and filter, and add a target node for the output location.
AnswerD

Glue Studio provides a visual job editor with drag-and-drop nodes for sources, transforms, and targets. A visual job can implement column renames and filters without writing code, and Glue generates the underlying PySpark script automatically. This directly meets the requirement to minimize development time and avoid manual PySpark coding, while still running on the Glue Spark engine for scale.

Why this answer

Glue Studio's visual job editor lets engineers build ETL pipelines by connecting source, transform, and target nodes without writing code. For a simple column rename and filter, this is the fastest path because Glue generates the Spark code and manages execution. The other options either require manual coding, use a different tool not integrated with Glue Studio jobs, or use query-based approaches that do not provide the same visual job experience.

Exam trap

The trap here is assuming that any visual data preparation tool, such as DataBrew, is interchangeable with Glue Studio's visual job editor for building ETL jobs.

1077
MCQeasy

A company needs to ingest real-time clickstream data from a web application into Amazon S3 for analytics. The data must be available within minutes of generation. Which AWS service should be used to capture and deliver this streaming data?

A.Amazon RDS
B.Amazon Kinesis Data Firehose
C.AWS Glue
D.Amazon Simple Queue Service (SQS)
AnswerB

Kinesis Data Firehose is a fully managed delivery stream that ingests streaming records and buffers them to S3 on a configurable interval, achieving near-real-time availability within minutes without consumer code. It matches the minutes-latency requirement for clickstream ingestion.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is a fully managed service designed to capture, transform, and load streaming data into Amazon S3, Redshift, Elasticsearch, or Splunk in near real-time (typically within 60 seconds). It directly addresses the requirement for ingesting real-time clickstream data and delivering it to S3 within minutes, without requiring custom code or manual scaling.

Exam trap

The trap here is confusing Amazon Kinesis Data Streams (which requires custom consumers and is not directly integrated with S3) with Amazon Kinesis Data Firehose (which is purpose-built for automated delivery to destinations like S3), leading candidates to overlook the 'within minutes' requirement and choose a service that needs additional components.

How to eliminate wrong answers

Option A (Amazon RDS) is wrong because it is a relational database service for transactional workloads, not designed for streaming data ingestion or direct delivery to S3; it would require additional ETL processes to move data to S3. Option C (AWS Glue) is wrong because it is a serverless ETL service for batch data processing and cataloging, not for real-time streaming capture; it can process data from S3 but does not ingest streaming data directly. Option D (Amazon Simple Queue Service (SQS)) is wrong because it is a message queue service for decoupling application components, not a streaming data delivery service; it does not automatically write data to S3 and requires custom consumers to do so.

1078
MCQmedium

A data pipeline using AWS Glue jobs is failing with 'Insufficient capacity' errors for Spark executors. Which action should the data engineer take to resolve this?

A.Reduce the number of workers in the Glue job configuration.
B.Increase the job timeout value.
C.Disable Spark UI logging.
D.Increase the number of workers (DPUs) in the Glue job configuration.
AnswerD

Increasing worker count directly addresses the 'Insufficient capacity' error, which occurs when Glue cannot provision enough DPUs to meet the job's requested executor count. Adding workers raises the total DPU allocation, allowing more Spark executors to launch and satisfy the pipeline's parallelism requirement.

Why this answer

The 'Insufficient capacity' error for Spark executors indicates that the Glue job is running out of resources (DPUs). Increasing the number of workers (DPUs) provides more compute capacity, allowing the job to allocate sufficient executors. Option A (reducing workers) would worsen the issue.

Option B (increasing timeout) does not add resources. Option C (disabling Spark UI) does not affect capacity. Therefore, increasing the number of workers (DPUs) is the correct resolution.

1079
Multi-Selecthard

A company is experiencing high costs from Amazon Redshift. The data engineer wants to optimize costs. Which THREE actions should the engineer take? (Choose THREE.)

Select 3 answers
A.Increase the frequency of automated snapshots.
B.Right-size the cluster based on workload analysis.
C.Increase the number of nodes to improve performance.
D.Purchase Reserved Instances for steady-state workloads.
E.Enable Concurrency Scaling and set up a usage limit.
AnswersB, D, E

Right-sizing matches provisioned cluster capacity to actual workload demand, eliminating spend on idle compute nodes. Redshift charges for running nodes regardless of utilisation, so analysing query concurrency, CPU and disk usage identifies over-provisioned clusters that can be resized to fewer or smaller nodes, directly reducing the cost driver described in the stem.

Why this answer

Option B is correct because right-sizing the cluster based on actual workload analysis ensures you are not paying for compute and storage capacity that the workload does not consume, directly reducing Redshift costs. Option D is correct because purchasing Reserved Instances for steady-state (predictable, always-on) workloads provides a significant discount over on-demand pricing for the same node usage. Option E is correct because Concurrency Scaling adds transient capacity only when queries queue, and pairing it with a usage limit caps how much extra concurrency-scaling compute can be billed, preventing runaway costs.

Option A is not correct because increasing automated snapshot frequency does not reduce cost and can actually increase storage charges for retained snapshots. Option C is not correct because adding more nodes increases both compute and storage cost, which is the opposite of cost optimization unless justified by a genuine performance need.

Exam trap

The trap here is that candidates confuse cost optimization with performance improvement, leading them to select 'Increase the number of nodes' (Option C) thinking it will reduce costs by improving efficiency, when in fact it increases costs.

1080
MCQmedium

A data engineer runs an AWS Glue ETL job that reads semi-structured JSON from Amazon S3, flattens nested arrays, and writes Parquet to a partitioned S3 location. Job runs are becoming expensive because Glue reprocesses all historical partitions on every run. The engineer wants subsequent runs to process only newly arrived data. Which approach should the engineer take with the LEAST operational overhead?

A.Schedule the job with a shorter interval so each run has less data to process.
B.Move the source data to an Amazon RDS table and use a Glue JDBC connection with a watermark column.
C.Create an AWS Glue job bookmark and enable it on the S3 data source, then let the job read only new objects on each run.
D.Add a WHERE clause to the DynamicFrame that filters records by the current date at runtime.
AnswerC

Job bookmarks persist state from prior runs and let the Glue job identify newly added S3 objects, so only incremental data is processed. This is the native, low-overhead mechanism for incremental processing in Glue ETL and directly solves the reprocessing cost problem without custom tracking.

Why this answer

Glue job bookmarks maintain run state and cause the reader to process only new or changed S3 objects since the last successful run. This eliminates repeated scanning of historical partitions with no custom code, unlike runtime filters or schedule changes, which still read all objects. Moving data to RDS is a much heavier architectural change than needed.

Exam trap

The trap here is assuming a date filter in the transform prevents historical data from being read, when filtering happens after the source objects are already listed and scanned.

1081
MCQhard

A data engineer runs an Apache Spark application on Amazon EMR that writes partitioned Parquet output to Amazon S3. A downstream AWS Glue crawler registers the table in the Data Catalog, but Athena queries return zero rows for partitions added by the most recent run, while older partitions query correctly. The S3 objects exist and are readable. Which action should the engineer take?

A.Convert the table to use Athena partition projection with a date-based template that covers the new prefix layout.
B.Run MSCK REPAIR TABLE or an ALTER TABLE ADD PARTITION statement against the table so the catalog discovers the newly written partitions.
C.Grant the crawler role s3:GetObject on the new partition prefixes so the crawler can inspect the files.
D.Increase the Athena query result reuse setting so previously cached partition metadata is invalidated.
AnswerB

Athena resolves partitions from the Data Catalog, not by listing S3 at query time. When Spark writes new partition prefixes without registering them, the catalog has no entries for those paths and the query planner filters them out, returning empty results for the new data. Registering the partitions with MSCK REPAIR TABLE or ALTER TABLE ADD PARTITION makes the new prefixes queryable.

Why this answer

Athena and Glue rely on partition metadata stored in the Data Catalog. Spark writes to S3 without automatically registering partitions unless the table is configured for it, so new prefixes remain invisible to the query planner and return no rows. Running MSCK REPAIR TABLE or adding partitions explicitly synchronizes the catalog with the S3 layout and makes the new data queryable.

Exam trap

The trap here is assuming Athena discovers partitions by listing Amazon S3 at query time, when it actually reads partition definitions from the Data Catalog.

1082
MCQhard

A data engineer is designing a streaming ingestion pipeline using Amazon Kinesis Data Streams. The stream has 10 shards, and the data volume is expected to grow by 50% over the next month. The engineer needs to ensure that the pipeline can scale without manual intervention. Which approach should be used?

A.Set up a CloudWatch Alarm to trigger a Lambda function to add shards
B.Use an Auto Scaling group to add more shards
C.Switch the Kinesis stream to on-demand capacity mode
D.Configure the stream to use a Lambda function that scales shards
AnswerC

On-demand capacity mode automatically scales shard throughput in response to traffic, removing the need to manually reshard as volume grows by 50%. It satisfies the no-manual-intervention constraint directly, unlike provisioned mode, which requires monitoring and explicit shard splits or merges to match demand.

Why this answer

Kinesis Data Streams on-demand capacity mode automatically scales shard capacity up or down based on observed throughput, eliminating the need to manually provision or add shards. It handles the 50% growth without any CloudWatch alarms, Lambda functions, or Auto Scaling groups. This is the only option that provides fully managed, hands-off scaling.

Exam trap

The trap here is confusing EC2 Auto Scaling concepts with Kinesis shard scaling, or assuming Lambda can programmatically add shards; the exam tests whether you know that only on-demand mode provides native automatic scaling for Kinesis Data Streams.

How to eliminate wrong answers

Option A is wrong because CloudWatch alarms plus Lambda shard addition is a custom, manual-provisioning workaround that requires you to manage scaling logic and shard counts yourself, not native auto-scaling. Option B is wrong because Auto Scaling groups apply to EC2 instances, not Kinesis shards; Kinesis does not integrate with ASGs. Option D is wrong because a Lambda function cannot natively scale Kinesis shards; shard scaling is controlled by the service (on-demand) or by explicit UpdateShardCount API calls, not by Lambda.

1083
MCQeasy

A company wants to migrate on-premises data to Amazon S3 using AWS DataSync. The data is stored on an NFS file server and the total volume is 50 TB. The network bandwidth between the on-premises data center and AWS is 1 Gbps (gigabit per second). What is the primary factor that will determine the total time required for the initial data transfer?

A.The available network bandwidth between on-premises and AWS
B.The number of S3 buckets used as the destination
C.The average file size in the dataset
D.The IOPS (I/O operations per second) of the on-premises NFS server
AnswerA

Available network bandwidth is the limiting factor: at 1 Gbps, roughly 125 MB/s theoretical throughput caps the transfer, so 50 TB needs many hours regardless of DataSync tuning. Disk read speed, S3 request rates and agent CPU rarely saturate a 1 Gbps link, making bandwidth the primary determinant of total elapsed time.

Why this answer

AWS DataSync transfers data over the network, so the primary constraint is the available bandwidth between the on-premises NFS server and AWS. With 50 TB of data and a 1 Gbps link, the theoretical minimum transfer time is approximately 50 TB * 8 / 1 Gbps = 400,000 seconds (~111 hours), but real-world throughput is lower due to protocol overhead, latency, and competing traffic. The network bandwidth directly dictates the maximum data transfer rate, making it the dominant factor for the initial transfer duration.

Exam trap

The trap here is that candidates may focus on the NFS server's IOPS or file size, assuming storage performance is the bottleneck, but the question explicitly provides a 1 Gbps bandwidth figure, signaling that network throughput is the key limiting factor for the initial transfer.

How to eliminate wrong answers

Option B is wrong because the number of S3 buckets does not affect transfer speed; DataSync can write to multiple buckets, but the throughput is still limited by the network pipe. Option C is wrong because while average file size can impact the number of file operations, DataSync uses parallel streams and can handle small files efficiently; the total data volume and bandwidth are the primary drivers, not file size. Option D is wrong because the NFS server's IOPS is rarely the bottleneck for a bulk transfer over a 1 Gbps link; DataSync reads files sequentially and the network bandwidth is typically the limiting factor, not the storage I/O performance.

1084
Multi-Selectmedium

A company uses Amazon S3 to store raw data and runs AWS Glue ETL jobs to transform it into Parquet. The data is then queried using Amazon Athena. Queries are slow and expensive due to high scan volumes. Which THREE design changes can improve query performance and reduce costs? (Select THREE.)

Select 3 answers
A.Increase the number of files by reducing file size to 1 MB
B.Convert the data to a columnar format like Parquet or ORC if not already
C.Compress the data using a splittable compression format like Snappy
D.Use bucketing on high-cardinality columns
E.Partition the data by commonly filtered columns such as date or region
AnswersB, C, E

Columnar formats such as Parquet or ORC store data by column, so Athena reads only the columns referenced in a query instead of every field. This directly cuts bytes scanned, which is the billed metric, lowering both query latency and cost for the existing Glue-produced data.

Why this answer

Option B is correct because columnar formats such as Parquet or ORC let Athena read only the columns referenced in a query, dramatically reducing the bytes scanned compared to row-based formats like CSV or JSON. Option C is correct because splittable compression such as Snappy (or Zstandard) reduces storage and scan volume while still allowing Athena to split large files across parallel readers, unlike non-splittable gzip on a single object. Option E is correct because partitioning by frequently filtered columns such as date or region enables partition pruning, so Athena skips reading irrelevant S3 prefixes and scans far less data.

Option A is not appropriate because shrinking files to 1 MB creates a huge number of small files, increasing S3 LIST and request overhead and hurting, not helping, Athena performance. Option D is not appropriate here because bucketing on high-cardinality columns does not reduce the data scanned by Athena and is generally used for join optimization in Hive/Spark, not for lowering Athena scan costs.

Exam trap

The trap here is that candidates may confuse bucketing with partitioning, or assume that increasing file count always improves parallelism, when in fact small files harm performance in distributed query engines like Athena.

1085
MCQeasy

A data engineer runs the command shown to check the encryption configuration of an S3 bucket. The output shows SSEAlgorithm: AES256. What does this mean?

A.The bucket uses SSE-S3 with Amazon S3-managed keys
B.The bucket uses SSE-KMS with a customer-managed key
C.The bucket uses SSE-C with customer-provided keys
D.The bucket does not have encryption enabled
AnswerA

SSEAlgorithm AES256 indicates server-side encryption with Amazon S3-managed keys (SSE-S3), where S3 owns and rotates the AES-256 keys. This satisfies the stem's observed output directly: SSE-KMS would report aws:kms, and SSE-C would not appear as a bucket default algorithm.

Why this answer

AES256 refers to SSE-S3, where Amazon S3 manages the encryption keys using AES-256. Option B (SSE-KMS) would show 'aws:kms'. Option C (SSE-C) would require the customer to provide keys.

Option D (no encryption) is incorrect because encryption is enabled.

1086
MCQmedium

A company is using Amazon Athena to query data in an S3 bucket. Queries are failing with the error 'HIVE_PATH_ALREADY_EXISTS'. The data is partitioned by year, month, day. What is the MOST likely cause?

A.A partition was manually added to the Glue Data Catalog that already exists
B.The data format in the partition is inconsistent with the table schema
C.The S3 location for the partition is empty
D.The IAM role used by Athena lacks s3:ListBucket permission on the bucket
AnswerA

Manually adding a partition to the AWS Glue Data Catalog when that partition already exists triggers HIVE_PATH_ALREADY_EXISTS, because Athena's metastore rejects duplicate partition locations for the year/month/day hierarchy. The error arises from the duplicate catalog entry, not from the underlying S3 data or query syntax.

Why this answer

The error 'HIVE_PATH_ALREADY_EXISTS' occurs in Athena when attempting to add a partition (via ALTER TABLE ADD PARTITION or MSCK REPAIR TABLE) that already exists in the Glue Data Catalog. Option B (inconsistent data format) would cause schema mismatch errors like 'HIVE_PARTITION_SCHEMA_MISMATCH', not this error. Option C (empty S3 location) would not cause this error; queries might succeed but return no results.

Option D (lack of s3:ListBucket permission) would cause permission errors like 'Access Denied'.

1087
MCQmedium

Refer to the exhibit. A data engineer queries AWS CloudTrail to investigate a PutObject event. What does the exhibit reveal about the object sensitive.csv?

A.The upload failed due to encryption mismatch.
B.The object was uploaded with server-side encryption using AWS KMS.
C.The object was not encrypted at rest.
D.The object was encrypted with SSE-S3.
AnswerB

The CloudTrail record shows the PutObject request carried the x-amz-server-side-encryption header set to aws:kms, confirming the object was written using SSE-KMS rather than SSE-S3 or SSE-C. This satisfies the investigation's need to identify the encryption mechanism applied to sensitive.csv at upload time.

Why this answer

The CloudTrail event contains `x-amz-server-side-encryption: aws:kms`, which confirms the object was uploaded with server-side encryption using AWS KMS (SSE-KMS). Option A is incorrect because the event shows a successful upload, not a failure. Option C is incorrect because the event indicates encryption was applied.

Option D is incorrect because SSE-S3 would show `AES256`, not `aws:kms`.

1088
Multi-Selectmedium

A data engineer is migrating a large Oracle data warehouse to Amazon Redshift. The engineer needs to ensure optimal performance. Which TWO practices should the engineer follow?

Select 2 answers
A.Choose appropriate sort keys based on common query patterns.
B.Design the schema as a normalized star schema with row-based storage.
C.Manually define compression encodings for each column.
D.Stage data in Amazon S3 before loading into Redshift.
E.Use DISTKEY to distribute data evenly across nodes.
AnswersA, E

Sort keys reduce the amount of data scanned.

Why this answer

Amazon Redshift uses sort keys to physically order data on disk, which allows the query optimizer to skip large blocks of data during scans via zone maps. Choosing sort keys based on common query patterns (e.g., range filters or frequent GROUP BY columns) dramatically reduces I/O and improves query performance, especially for large tables.

Exam trap

The trap here is that candidates often confuse Redshift's columnar storage with row-based storage and assume a normalized star schema is optimal, when in fact Redshift is designed for denormalized, columnar tables with explicit sort and distribution keys.

1089
MCQmedium

A data engineer is building an AWS Glue ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3. The DynamoDB table has a large number of items, and the engineer needs to ensure the job reads the data efficiently without consuming too much provisioned throughput. Which method should the engineer use to read from DynamoDB?

A.Use Amazon Kinesis Data Streams to stream changes from DynamoDB to S3, then read from S3.
B.Use AWS Glue's native DynamoDB connector with the read throughput parameter set to a high value.
C.Use AWS Glue's DynamoDB connector with the 'dynamodb.scan' option to scan the table.
D.Use the DynamoDB export to S3 feature to export the table to S3, then read from S3 in the Glue job.
AnswerD

DynamoDB export to S3 is a serverless, efficient way to export large tables without consuming read capacity. The export creates a consistent snapshot in S3, which Glue can then read in parallel. This method avoids impacting the table's provisioned throughput and is recommended for large-scale reads.

Why this answer

DynamoDB export to S3 allows you to export a full table snapshot to S3 without consuming read capacity, making it ideal for large-scale ETL reads. The Glue job can then read the exported data from S3 in parallel, achieving high throughput without impacting the DynamoDB table's performance.

Exam trap

The trap here is assuming that using the native DynamoDB connector with high throughput is the best way, when it can actually degrade performance by consuming provisioned capacity.

1090
MCQmedium

A data engineer is configuring an S3 bucket for storing sensitive customer data. The bucket must be encrypted at rest using an AWS Key Management Service (KMS) key that is managed by the data engineering team. The team wants to ensure that only users with explicit permission can decrypt the data. Which S3 encryption option should be used?

A.SSE-KMS
B.Client-side encryption
C.SSE-S3
D.SSE-C
AnswerA

SSE-KMS encrypts objects with a KMS key, and decryption requires kms:Decrypt permission on that key. This satisfies the requirement for team-managed keys and explicit decrypt authorisation, unlike SSE-S3 where AWS owns the key and access is governed solely by S3 permissions.

Why this answer

SSE-KMS is the correct option because it uses a customer-managed AWS KMS key, allowing the data engineering team to control access and permissions for decryption. Client-side encryption is not an S3 server-side encryption option and does not use KMS. SSE-S3 uses Amazon S3-managed keys, which do not provide customer-controlled access.

SSE-C requires the customer to manage their own encryption keys and does not use KMS, nor does it allow the same level of access control as a CMK.

1091
MCQeasy

A data engineer needs to ingest log files from multiple EC2 instances into Amazon S3. The logs are written to local disk on each instance. The engineer wants a simple agent-based solution that can collect, compress, and upload logs to S3 with minimal configuration. The solution must support incremental uploads (only new log lines) and handle log rotation. What should the engineer use?

A.Install and configure Amazon CloudWatch Agent to collect logs and send them to Amazon CloudWatch Logs, then use a subscription filter to export logs to S3.
B.Use AWS CLI cp command with --recursive in a cron job to copy logs to S3 every minute.
C.Install AWS DataSync agent on each EC2 instance to sync logs to S3 daily.
D.Use an S3 sync command from the AWS CLI scheduled every hour.
AnswerA

Kinesis Agent tails log files, compresses, and sends to CloudWatch Logs; export to S3 can be automated.

Why this answer

Amazon CloudWatch Agent is a lightweight agent that can tail log files, compress them on the fly, and send them to CloudWatch Logs. From CloudWatch Logs, a subscription filter can export the logs to Amazon S3, supporting incremental uploads and log rotation. Option B (AWS CLI cp) is manual and does not handle incremental uploads efficiently.

Option C (AWS DataSync) is designed for bulk data transfers, not real-time log ingestion. Option D (S3 sync) is also not real-time and lacks agent-based tailing and compression.

1092
MCQeasy

A retail company uses Amazon DynamoDB to store product catalog data. The table has a partition key of ProductID and a sort key of Category. The company needs to retrieve all products in a specific category, sorted by ProductID. Which operation should be used?

A.Create a global secondary index (GSI) with Category as the partition key and ProductID as the sort key, then query the GSI.
B.Scan the table with a filter expression on Category.
C.Query the table using the partition key and sort key condition.
D.Use the BatchGetItem operation with Category as a key.
AnswerA

A GSI allows querying on non-primary key attributes. By setting Category as the partition key and ProductID as the sort key, you can efficiently query all products in a category, sorted by ProductID. This is the optimal solution for this access pattern.

Why this answer

To efficiently retrieve all products in a specific category sorted by ProductID, a global secondary index (GSI) is needed. The GSI should use Category as the partition key and ProductID as the sort key. This allows Query operations on the GSI to return the desired items in sorted order, leveraging DynamoDB's indexing capabilities for performance and cost efficiency.

Exam trap

The trap here is assuming that a Scan with a filter or a Query on the base table can satisfy the access pattern, when in fact the base table's key schema does not support querying by Category with sorting by ProductID.

1093
MCQmedium

A data engineer is using AWS Glue DataBrew to profile a dataset stored in Amazon S3. The profile shows that a column named country contains values such as 'US', 'usa', 'United States', and 'U.S.A.' The engineer needs to standardize these values to a single canonical form before loading the data into Amazon Redshift. Which DataBrew transformation should the engineer apply?

A.Use the 'Remove special characters' transform to strip punctuation.
B.Use the 'Replace value or pattern' transform to replace each variant with 'United States'.
C.Use the 'Change case' transform to convert all values to uppercase.
D.Use the 'Flag duplicate values' transform to identify repeated entries.
AnswerB

The 'Replace value or pattern' transform in DataBrew lets the engineer map each observed variant to a canonical value, handling 'US', 'usa', 'United States', and 'U.S.A.' in one recipe step. It matches the profiling output directly and produces a deterministic result, which is exactly what standardization requires before loading into Redshift.

Why this answer

Standardizing a column with several representations of the same entity requires explicit value mapping, which the 'Replace value or pattern' transform provides. Case conversion, punctuation removal, and duplicate flagging each address only part of the problem or change nothing in the stored values. The replace transform is the only option that yields one canonical value for all variants.

Exam trap

The trap here is treating case normalization or punctuation cleanup as sufficient standardization when the column contains entirely different spellings of the same entity.

1094
MCQmedium

A data engineer is designing a data lake on S3 with sensitive data. The security policy mandates that data must be encrypted at rest and in transit, and that an inventory of all objects must be maintained for compliance. Which actions should be taken?

A.Enforce HTTPS via bucket policy, enable default SSE-S3 encryption, and enable S3 Inventory.
B.Use SSE-KMS encryption and enable CloudTrail for S3 events.
C.Enable S3 default encryption using SSE-S3 and enable S3 Inventory.
D.Enforce HTTPS using bucket policy and enable S3 Server Access Logging.
AnswerA

Enforcing HTTPS via bucket policy satisfies the in-transit encryption mandate, while default SSE-S3 provides AES-256 server-side encryption at rest without key management overhead. S3 Inventory delivers scheduled CSV or Parquet reports of all objects and their metadata, meeting the compliance inventory requirement. Together these three controls cover every stated constraint.

Why this answer

It covers all requirements: encryption in transit (HTTPS enforcement via bucket policy), encryption at rest (default SSE-S3), and compliance inventory (S3 Inventory). Option B uses SSE-KMS which is not required and lacks inventory. Option C includes at-rest encryption and inventory but misses in-transit encryption.

Option D includes in-transit encryption but lacks at-rest encryption and compliance inventory.

1095
Multi-Selectmedium

Which THREE storage classes in Amazon S3 are designed for infrequently accessed data with millisecond retrieval times? (Select THREE.)

Select 3 answers
A.S3 Glacier Flexible Retrieval
B.S3 One Zone-IA
C.S3 Glacier Deep Archive
D.S3 Intelligent-Tiering
E.S3 Standard-IA
AnswersB, D, E

S3 One Zone-IA stores data in a single Availability Zone, delivering millisecond retrieval while suiting infrequently accessed data. It satisfies the stem's latency constraint because retrieval remains immediate, unlike Glacier tiers requiring minutes to hours. Lower durability than Standard-IA is the trade-off, but the millisecond requirement is met.

Why this answer

S3 One Zone-IA (B) is correct because it is an infrequent-access class that stores data in a single Availability Zone while still providing millisecond retrieval latency. S3 Intelligent-Tiering (D) is correct because it automatically moves objects between access tiers based on changing access patterns, and its Frequent and Infrequent Access tiers both deliver millisecond retrieval times. S3 Standard-IA (E) is correct because it is explicitly designed for infrequently accessed data and provides the same low-latency, millisecond retrieval performance as S3 Standard.

S3 Glacier Flexible Retrieval (A) is not correct here because its retrieval options range from minutes to hours, not milliseconds, and S3 Glacier Deep Archive (C) is not correct because it is the lowest-cost archive class with retrieval times typically within 12 hours.

Exam trap

The trap here is that candidates often confuse S3 Glacier Flexible Retrieval or S3 Glacier Deep Archive as having millisecond retrieval times, but these classes are designed for archival access with retrieval times measured in minutes or hours, not milliseconds.

1096
MCQmedium

A company is using Amazon RDS for MySQL with Multi-AZ deployment. The primary DB instance experiences a hardware failure, causing automatic failover to the standby. After the failover, the application reports that the database endpoint is unreachable for about 60 seconds. What is the MOST likely cause?

A.The standby instance took longer than expected to promote to primary.
B.The standby instance was not in a synchronized state and required a manual promotion.
C.The application was using the wrong endpoint and needed to be reconfigured.
D.The DNS record for the DB instance endpoint needed to update to point to the new primary.
AnswerD

Multi-AZ failover promotes the standby and repoints the DB instance's DNS CNAME to the new primary; clients caching the old record cannot connect until the TTL expires, producing the roughly 60-second outage described. This satisfies the stem's constraint that the endpoint itself was unreachable, not merely slow.

Why this answer

After an automatic failover in Amazon RDS Multi-AZ, the DNS record for the DB instance endpoint is updated to point to the new primary. This DNS change can take up to 60 seconds to propagate, during which the application may receive 'unreachable' errors if it caches the old DNS resolution. The 60-second outage aligns with the typical TTL (Time To Live) of 30 seconds for RDS DNS records plus propagation delays.

Exam trap

The trap here is that candidates assume the standby promotion itself causes the delay, but AWS specifically designs the promotion to be fast, and the real bottleneck is DNS propagation and client caching.

How to eliminate wrong answers

Option A is wrong because the standby promotion itself is nearly instantaneous in RDS Multi-AZ; the delay is not due to promotion time but DNS propagation. Option B is wrong because RDS Multi-AZ automatically synchronizes the standby synchronously, and no manual promotion is required—the failover is fully automated. Option C is wrong because the application uses the same RDS endpoint (CNAME) before and after failover; no reconfiguration is needed.

1097
MCQmedium

A data engineer maintains an AWS Glue ETL job that processes JSON files from Amazon S3 and writes Parquet to another S3 location. The job has been running successfully for months. Recently, the job started failing intermittently with the error 'Unable to infer schema for JSON'. The engineer confirms the source bucket contains valid JSON files. Which action should the engineer take to resolve the failure?

A.Enable job bookmarks to track previously processed files and avoid reprocessing.
B.Increase the number of AWS Glue DPUs allocated to the job to handle larger file sizes.
C.Define an explicit schema using a Glue Data Catalog table or a DynamicFrame with a specified schema, instead of relying on schema inference.
D.Convert the source JSON files to CSV format before running the job, then update the job to read CSV.
AnswerC

When JSON files have inconsistent structures, Glue's automatic schema inference can fail. Providing an explicit schema via a Data Catalog table or by passing a schema to the DynamicFrame resolves the ambiguity. This ensures the job can parse the data reliably regardless of variations, and is the recommended approach for production workloads with evolving or irregular JSON.

Why this answer

The error 'Unable to infer schema for JSON' occurs when AWS Glue cannot automatically determine a consistent schema from the source files. This often happens when JSON records have varying fields or data types. Defining an explicit schema through the Glue Data Catalog or programmatically ensures the job can parse the data without relying on inference, resolving the failure while maintaining data integrity.

Exam trap

The trap here is assuming that increasing compute resources or changing file formats will fix schema inference errors, when the real solution is to provide an explicit schema.

1098
MCQeasy

A data engineer needs to ingest streaming data from an IoT fleet into Amazon S3 for near-real-time analytics. The data volume is approximately 5 GB per hour, and each event is less than 1 KB. Which AWS service should be used as the ingestion endpoint?

A.AWS IoT Core
B.AWS DataSync
C.Amazon AppFlow
D.Amazon Kinesis Data Streams
AnswerA

AWS IoT Core ingests high-volume, small-payload device telemetry and can route it directly to Amazon S3 via rules, meeting the 5 GB per hour near-real-time requirement. Its native MQTT support suits sub-1 KB events from a fleet, unlike services designed for batch or large-object transfer.

Why this answer

AWS IoT Core is purpose-built for ingesting data from IoT devices, supporting MQTT, HTTP, and WebSocket protocols. It can handle millions of devices and high-throughput, small-message payloads (each event <1 KB) and integrates directly with Amazon S3 via IoT Core rules, making it the ideal ingestion endpoint for near-real-time analytics on streaming IoT data.

Exam trap

The trap here is that candidates often default to Amazon Kinesis Data Streams for any streaming workload, overlooking that AWS IoT Core is the specialized, fully managed service designed specifically for IoT device ingestion, with native MQTT support and direct S3 integration via rules.

How to eliminate wrong answers

Option B (AWS DataSync) is wrong because it is designed for one-time or scheduled bulk data transfers between on-premises storage and AWS, not for continuous, near-real-time streaming ingestion from IoT devices. Option C (Amazon AppFlow) is wrong because it is a fully managed integration service for transferring data between SaaS applications (e.g., Salesforce, Slack) and AWS, not for ingesting IoT device telemetry streams. Option D (Amazon Kinesis Data Streams) is wrong because while it can ingest streaming data, it is a generic stream processing service that requires additional configuration (e.g., Kinesis Data Firehose) to write to S3, and it is not the dedicated IoT ingestion endpoint; AWS IoT Core is the recommended first-hop for IoT data.

1099
MCQeasy

A data engineer needs to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real time, and the engineer wants to minimize operational overhead by using a fully managed service that can also transform the data format from JSON to Parquet. Which AWS service should be used?

A.AWS Glue ETL job
B.Amazon Kinesis Data Firehose
C.Amazon Kinesis Client Library (KCL) application
D.AWS Database Migration Service (AWS DMS)
AnswerB

Amazon Kinesis Data Firehose is a fully managed service that can read from Kinesis Data Streams, transform data using AWS Lambda or built-in format conversion, and deliver to Amazon S3. It supports converting JSON to Parquet using the built-in record format conversion feature, which requires a AWS Glue table schema. This minimizes operational overhead and provides near real-time delivery.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is fully managed, can ingest from Kinesis Data Streams, and supports built-in conversion of JSON to Parquet before delivering to Amazon S3. It requires minimal operational effort and provides near real-time delivery. The other options either require custom code, are batch-oriented, or do not support the required source and transformations.

Exam trap

The trap here is confusing Kinesis Data Firehose with Kinesis Data Streams or Kinesis Client Library, or assuming AWS Glue can handle real-time ingestion with format conversion out of the box.

1100
MCQeasy

A company has an S3 bucket that stores logs for compliance. The compliance team requires that objects are retained for 7 years and cannot be deleted or overwritten. Which S3 feature should be used?

A.Enable S3 Object Lock with retention mode COMPLIANCE and a retention period of 7 years
B.Enable MFA Delete on the bucket
C.Configure an S3 bucket policy that denies delete and overwrite actions
D.Enable S3 Versioning and configure a lifecycle policy to expire objects after 7 years
AnswerA

S3 Object Lock in COMPLIANCE mode enforces a WORM retention period that no user, including the root account, can shorten or bypass, directly satisfying the requirement that objects cannot be deleted or overwritten for 7 years. GOVERNANCE mode would permit privileged users to alter retention, so it fails this constraint.

Why this answer

S3 Object Lock in COMPLIANCE mode enforces a WORM (Write Once Read Many) model where no user, including the root account, can delete or overwrite the object version until the retention period expires. Setting a 7-year retention period directly satisfies the compliance requirement that objects cannot be deleted or overwritten for 7 years. COMPLIANCE mode is the strictest retention mode and is specifically designed for regulatory retention scenarios.

Exam trap

DEA-C01 often tests the misconception that MFA Delete or bucket policies provide immutability — candidates confuse access control with WORM retention, but only Object Lock in COMPLIANCE mode guarantees non-deletable, non-overwritable objects.

How to eliminate wrong answers

Option B is wrong because MFA Delete only requires additional authentication for delete operations; it does not prevent deletion or overwriting, and a user with MFA can still delete objects. Option C is wrong because a bucket policy denying delete and overwrite actions can be modified or removed by an administrator, so it is not tamper-proof and does not provide immutable retention. Option D is wrong because versioning plus a lifecycle expiration policy allows objects to be deleted after 7 years but does not prevent deletion or overwriting before 7 years; lifecycle rules can also be changed at any time.

1101
MCQeasy

A data engineer needs to store large amounts of data that is accessed infrequently but must be retrieved immediately when needed. Which Amazon S3 storage class is most cost-effective?

A.S3 Intelligent-Tiering
B.S3 One Zone-IA
C.S3 Standard-IA
D.S3 Glacier Deep Archive
AnswerC

S3 Standard-IA is designed for infrequent access with millisecond retrieval.

Why this answer

S3 Standard-IA (Infrequent Access) is the most cost-effective choice because it offers low per-GB storage costs for data accessed infrequently, while still providing millisecond retrieval latency for immediate access when needed. This matches the requirement of storing large amounts of data that is rarely accessed but must be available instantly.

Exam trap

The DEA-C01 exam often tests the misconception that S3 One Zone-IA is a cheaper alternative for infrequent access, but the trap is that it sacrifices durability by storing data in a single Availability Zone, which is not suitable for data that must be reliably retrieved immediately.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering automatically moves data between access tiers based on usage patterns, but it incurs a monthly monitoring and automation fee per object, making it less cost-effective for purely infrequent access patterns with no variable usage. Option B is wrong because S3 One Zone-IA stores data in a single Availability Zone, which risks data loss if that AZ fails, and it does not meet the implied durability requirement for data that must be retrievable immediately. Option D is wrong because S3 Glacier Deep Archive is designed for archival data with retrieval times of 12 to 48 hours, not immediate retrieval, and thus fails the 'retrieved immediately' requirement.

1102
MCQeasy

A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3. The data is stored as CSV files. The downstream team requires the data to be in Apache Parquet format. Which change should the data engineer make to the DMS task?

A.Modify the DMS task to use Apache Parquet as the target table preparation mode.
B.Add an S3 lifecycle rule to convert CSV to Parquet.
C.Change the DMS task to use full load instead of continuous replication.
D.Configure a Lambda function to transform data after DMS writes to S3.
AnswerA

DMS can write directly in Parquet format.

Why this answer

AWS DMS supports specifying Apache Parquet as the target data format for Amazon S3 targets directly within the task configuration. By setting the 'Data format' to 'Parquet' in the S3 target endpoint or task settings, DMS automatically converts the replicated data into Parquet files, eliminating the need for post-processing. This is the most efficient and native approach to meet the downstream team's requirement.

Exam trap

The trap here is that candidates may assume DMS only supports CSV for S3 targets, overlooking the built-in Parquet option, and instead choose a complex workaround like Lambda or lifecycle rules.

How to eliminate wrong answers

Option B is wrong because S3 lifecycle rules can transition objects between storage classes or expire them, but they cannot convert file formats (e.g., CSV to Parquet). Option C is wrong because changing from continuous replication to full load would stop ongoing data synchronization and does not address the format conversion requirement. Option D is wrong because while a Lambda function could transform CSV to Parquet, it introduces unnecessary complexity, latency, and cost compared to the native DMS capability; DMS can directly write Parquet without additional services.

1103
MCQeasy

A data engineer is building an AWS Glue job that reads from a JDBC source and must retrieve the database password at runtime without hardcoding it in the script or job parameters in plaintext. The company already stores the password in AWS Secrets Manager. Which action should the engineer take?

A.Grant the Glue job role secretsmanager:GetSecretValue on the secret ARN and retrieve the secret in the job using the Glue Secrets Manager connection property.
B.Embed the password directly in the Glue ETL script and restrict access to the script in Amazon S3.
C.Use an IAM database authentication token for the JDBC connection instead of a password.
D.Store the password in an AWS Glue job parameter and mark it as encrypted using a KMS key.
AnswerA

AWS Glue integrates with Secrets Manager through the connection's secretId property, letting the job fetch credentials at runtime. Granting secretsmanager:GetSecretValue scoped to the specific secret ARN provides least-privilege access. This avoids embedding plaintext credentials in scripts or job parameters while giving the job the password it needs.

Why this answer

AWS Glue connections support a secretId property that fetches credentials from Secrets Manager at runtime. Granting the job role GetSecretValue on the specific secret keeps access narrow and avoids plaintext credentials in scripts or parameters, while also enabling rotation and audit through Secrets Manager.

Exam trap

The trap here is assuming encrypted job parameters are safe, when the real goal is to keep credentials out of the job definition entirely via a secrets store.

1104
Multi-Selecthard

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 source with many small JSON files and writes to Amazon S3 in Parquet. The job runs slowly and produces many tiny output files. The engineer wants to improve throughput and reduce the number of output files without changing the source data layout. (Choose two.)

Select 2 answers
A.Call coalesce or repartition on the DataFrame before writing the Parquet output.
B.Enable the AWS Glue groupFiles option to coalesce input files during the read.
C.Increase the number of Glue DPUs allocated to the job.
D.Enable AWS Glue job bookmarks to skip previously processed files.
E.Convert the output format from Parquet to CSV to reduce file count.
AnswersA, B

Repartitioning or coalescing the DataFrame controls how many partitions exist at write time, and each partition produces one output file. Reducing partitions before the write consolidates output into fewer, larger Parquet files, which improves downstream query performance and directly addresses the many-tiny-output-files problem.

Why this answer

Small input files cause per-object overhead that slows reads, and the output file count equals the number of write-time partitions. Grouping input files with the groupFiles option consolidates reads, while repartitioning or coalescing the DataFrame before writing reduces the number of output partitions. Together these improve throughput and produce fewer, larger Parquet files without altering the source layout.

Exam trap

The trap here is assuming that adding more DPUs will consolidate output files, when output file count is determined by write-time partitions, not by worker capacity.

1105
Multi-Selectmedium

A company is using Amazon DynamoDB as a data store for a real-time application. The application reads a single item by primary key and occasionally updates it. The data engineer notices high read latency during peak hours. Which TWO actions would most effectively reduce read latency?

Select 2 answers
A.Increase the read capacity units for the table.
B.Enable DynamoDB global tables.
C.Add a local secondary index on the table.
D.Disable auto-scaling and set a fixed read capacity.
E.Enable DynamoDB Accelerator (DAX) for the table.
AnswersA, E

Increasing read capacity units raises the table's provisioned throughput ceiling, so reads are not throttled when consumed capacity approaches the limit during peak hours. For a workload dominated by single-item GetItem calls, this directly addresses the constraint causing latency: insufficient provisioned read throughput for the peak request rate.

Why this answer

Option A is correct because the workload is read-heavy with single-item reads by primary key, and if the table is provisioned with insufficient read capacity units (RCUs), requests will throttle or queue during peak hours, so raising the provisioned RCUs directly increases available read throughput and reduces latency. Option E is correct because DynamoDB Accelerator (DAX) is an in-memory cache purpose-built for DynamoDB that serves eventually consistent reads of individual items in microseconds, which is ideal for this read-by-primary-key pattern and offloads repeated reads from the table. Option B is not appropriate because global tables provide multi-Region active-active replication for availability and locality, not lower latency for a single-Region read pattern, and they add replication overhead.

Option C is not appropriate because a local secondary index only provides an alternative sort key on the same partition key and does not speed up reads that already use the primary key. Option D is wrong because disabling auto-scaling and fixing read capacity removes the ability to scale with peak demand and would likely worsen throttling and latency.

Exam trap

The trap here is that candidates may confuse global tables or secondary indexes as solutions for read latency, when in fact they address different concerns (disaster recovery and query flexibility), while the correct approach is to either increase provisioned throughput or implement a caching layer like DAX.

1106
MCQeasy

A company uses Amazon RDS for PostgreSQL to store customer data. The data engineer needs to ensure that the database can be restored to any point in time within the last 35 days. The engineer also wants to minimize the impact on the production database during backups. What should the engineer do?

A.Enable automated backups and set the backup retention period to 35 days.
B.Create a manual snapshot every day and retain them for 35 days.
C.Enable Multi-AZ deployment and rely on the standby replica for backups.
D.Configure a read replica and take snapshots from the read replica.
AnswerA

Amazon RDS automated backups enable point-in-time recovery to any second within the retention period, up to 35 days. They are taken during a daily backup window and continuously archive transaction logs to S3, with minimal impact on the production database. Setting the retention period to 35 days directly meets the requirement.

Why this answer

Amazon RDS automated backups provide point-in-time recovery to any second within the retention period, which can be set up to 35 days. They are managed by RDS, require no manual intervention, and have minimal performance impact. Manual snapshots, Multi-AZ, and read replicas do not offer the same continuous PITR capability.

Exam trap

The trap here is assuming that Multi-AZ or read replicas automatically provide point-in-time recovery, when only automated backups enable PITR within the retention period.

1107
MCQmedium

A data engineer maintains an Amazon Kinesis Data Streams pipeline that feeds an AWS Lambda consumer. During traffic spikes, the Lambda function is throttled and records are reprocessed, causing duplicate entries in the downstream Amazon S3 sink. The engineer needs to reduce duplicates with the LEAST code change. What should the engineer do?

A.Configure the event source mapping's StartingPosition to LATEST and enable enhanced fan-out on the stream.
B.Increase the Lambda function's reserved concurrency and set the event source mapping's MaximumRetryAttempts to a low value.
C.Enable the 'ReportBatchItemFailures' feature on the event source mapping and return the failed sequence numbers so only failed records are retried.
D.Make the Lambda function idempotent by using the Kinesis sequence number as a deduplication key before writing to S3.
AnswerD

Kinesis records carry a unique sequence number per shard, so using it as an idempotency key lets the Lambda function skip records it has already processed. This directly addresses duplicate writes at the sink with minimal code change, and it remains effective regardless of retries or throttling because reprocessed records are recognized and discarded.

Why this answer

Duplicate delivery in Kinesis-to-Lambda pipelines is expected because Lambda retries batches after throttling or errors, and Kinesis guarantees at-least-once delivery. Making the consumer idempotent by keying on the Kinesis sequence number ensures that reprocessed records are detected and not written twice, which requires only a small change inside the function.

Exam trap

The trap here is treating throughput tuning such as concurrency or enhanced fan-out as a fix for duplicates, when duplicates are an at-least-once delivery property that only idempotency resolves.

1108
MCQmedium

A company is designing a data lake on AWS and must comply with GDPR requirements. The company needs to implement data masking for personally identifiable information (PII) columns in Amazon Redshift. Which feature should be used?

A.Use Amazon RDS Proxy to intercept queries
B.Amazon S3 Object Lambda to mask data on the fly
C.Create views in Redshift that apply masking functions
D.AWS Lake Formation row-level security
AnswerC

Dynamic data masking in Amazon Redshift applies masking policies at query time, so PII columns return redacted values without altering stored data. Attaching these policies to roles satisfies GDPR's data-minimisation requirement, and unlike views, masking persists across all queries against the table, including ad-hoc analyst access.

Why this answer

Amazon Redshift supports dynamic data masking through views that apply masking functions, such as using CASE statements or custom masking functions to obfuscate PII columns. Option A is incorrect because Amazon RDS Proxy is a connection proxy for RDS databases and does not provide data masking capabilities for Redshift. Option B is incorrect because Amazon S3 Object Lambda is used to transform data in S3, not to mask data in Redshift queries.

Option D is incorrect because AWS Lake Formation row-level security filters rows based on permissions but does not mask or obfuscate column values; it is for access control, not data masking.

1109
MCQmedium

A data engineer is troubleshooting a step function that orchestrates ETL jobs. The state machine fails with 'State Machine Execution Throttled' error. What should the engineer do to resolve this?

A.Reduce the number of steps in the state machine.
B.Set up a CloudWatch alarm to detect throttling and retry.
C.Adjust the API rate limits in the state machine definition.
D.Request a service quota increase for concurrent executions.
AnswerD

Requesting a quota increase for concurrent executions directly addresses the throttling constraint: Step Functions limits how many state machine executions can run simultaneously per account and Region. When ETL jobs peak and hit that ceiling, new executions are rejected with 'Execution Throttled'. Raising the quota via Service Quotas lifts the ceiling, allowing the required parallelism.

Why this answer

The 'State Machine Execution Throttled' error in AWS Step Functions occurs when the account hits the service quota for concurrent state machine executions (default 1,000 per region, or lower for certain execution types). The correct remediation is to request a service quota increase via the AWS Service Quotas console for the 'Concurrent executions' quota of Step Functions. This directly addresses the throttling cause rather than masking it.

Exam trap

DEA-C01 often tests the misconception that throttling errors are code or configuration bugs to be fixed inside the state machine, when in reality they are service quota limits that require a quota increase request.

How to eliminate wrong answers

Option A is wrong because reducing the number of steps in a state machine does not affect the concurrent execution quota — throttling is about how many executions run simultaneously, not how many states each execution contains. Option B is wrong because a CloudWatch alarm only detects and notifies about throttling; it does not increase capacity, and retrying without more quota will simply fail again. Option C is wrong because API rate limits for Step Functions are AWS-managed service quotas, not configurable parameters inside the state machine's Amazon States Language (ASL) definition.

1110
MCQmedium

A company uses AWS Kinesis Data Streams to ingest real-time data. The data engineer notices that the stream's 'WriteProvisionedThroughputExceeded' error occurs frequently during peaks. Which action should be taken to resolve this issue?

A.Increase the number of shards in the stream.
B.Modify the producer to use a different partition key.
C.Compress the data before sending to the stream.
D.Enable enhanced fan-out for consumers.
AnswerA

Increasing shard count raises the stream's total write capacity, since each shard provides a fixed 1 MB/s and 1,000 records/s ingest limit. The WriteProvisionedThroughputExceeded error signals that producers are exceeding the aggregate provisioned throughput, so adding shards directly satisfies the peak-demand constraint described in the stem.

Why this answer

The 'WriteProvisionedThroughputExceeded' error occurs when the stream's write capacity is exceeded. Increasing the number of shards (A) increases the stream's write capacity, as each shard provides a fixed write throughput (1 MB/s or 1000 records/s). This directly resolves the issue by adding more capacity to handle peak loads.

Exam trap

DEA-C01 often tests the misconception that changing the partition key or compressing data can resolve throughput errors, but the fundamental solution is to increase shard capacity.

How to eliminate wrong answers

Option B is wrong because changing the partition key may distribute data more evenly but does not increase the total write capacity; if the overall throughput exceeds capacity, the error persists. Option C is wrong because compressing data reduces the size of each record, which can help if the bottleneck is record size, but it does not increase the number of records per second or the total throughput capacity. Option D is wrong because enhanced fan-out is for consumers, improving read throughput, not write capacity.

1111
MCQeasy

A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift table on a daily schedule. The data is in CSV format and the schema matches. Which service is simplest for this batch ingestion?

A.Amazon Redshift COPY command
B.AWS Glue ETL job with JDBC connection
C.AWS Data Pipeline
D.Amazon Athena CREATE TABLE AS SELECT
AnswerA

The Redshift COPY command loads CSV files directly from Amazon S3 into a table in parallel, parsing and distributing data across slices natively. It requires no cluster provisioning or transformation code, making it the simplest option for scheduled batch ingestion with matching schemas.

Why this answer

The Amazon Redshift COPY command is the simplest and most efficient method for batch loading data from Amazon S3 into Redshift when the schema matches and the data is in CSV format. It leverages Redshift's massively parallel processing (MPP) architecture to read data directly from S3, automatically handling compression, encryption, and error logging without requiring any intermediate services or custom code.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing AWS Glue or Data Pipeline, forgetting that Redshift's native COPY command is purpose-built for high-speed, parallel batch ingestion from S3 with minimal configuration.

How to eliminate wrong answers

Option B is wrong because AWS Glue ETL with JDBC connection introduces unnecessary complexity and overhead for a simple schema-matching CSV load; Glue is better suited for complex transformations or semi-structured data, not for a direct COPY operation. Option C is wrong because AWS Data Pipeline is a legacy orchestration service that requires defining pipelines, schedules, and activities, adding operational overhead compared to the single COPY command. Option D is wrong because Amazon Athena CREATE TABLE AS SELECT (CTAS) writes query results to a new table in S3, not into Redshift; it cannot directly ingest data into a Redshift table.

1112
MCQhard

A data engineer is troubleshooting an issue where an IAM role used by AWS Glue cannot read data from an S3 bucket encrypted with SSE-KMS. The bucket policy allows the role to perform s3:GetObject. What additional permission is needed?

A.s3:GetObjectVersion
B.kms:Decrypt on the KMS key
C.s3:GetObjectAcl
D.kms:GenerateDataKey on the KMS key
AnswerB

SSE-KMS encrypts objects with a KMS key, so s3:GetObject alone is insufficient: S3 must call KMS to unwrap the data key on the caller's behalf. The role therefore needs kms:Decrypt on that key, satisfying the stem's requirement to read the SSE-KMS-encrypted objects.

Why this answer

For SSE-KMS, the IAM role needs kms:Decrypt permission on the KMS key to read encrypted objects. Option A (s3:GetObjectVersion) is not required because the bucket policy already allows s3:GetObject; versioning is not relevant here. Option C (s3:GetObjectAcl) is for access control lists, not encryption.

Option D (kms:GenerateDataKey) is used for encrypting new objects, not reading existing ones. Therefore, the correct answer is B.

1113
MCQmedium

A data engineer needs to transfer 50 TB of historical data from an on-premises HDFS cluster to Amazon S3. The on-premises network has a 1 Gbps link to AWS. The transfer must complete within 5 days. Which solution is MOST cost-effective and meets the requirements?

A.Use Amazon S3 Transfer Acceleration to speed up the transfer over the internet.
B.Use AWS DataSync to transfer the data over the existing network link.
C.Use AWS Snowball Edge to physically transfer the data.
D.Use AWS Direct Connect to establish a dedicated network connection.
AnswerC

AWS Snowball Edge is a physical device used for offline data transfer. It can handle large volumes faster than network transfer and is cost-effective for this scenario.

Why this answer

AWS Snowball Edge is a physical device that can transfer large amounts of data faster than a network link, especially with a 1 Gbps link that would take about 4.6 days for 50 TB (theoretical max, but actual throughput will be lower due to overhead). Snowball Edge can transfer 50 TB in a few days and is cost-effective for large data volumes. Option A (Amazon S3 Transfer Acceleration) speeds up transfers but still limited by network bandwidth.

Option B (AWS DataSync) is efficient for online transfers but may not meet the 5-day deadline over 1 Gbps. Option D (AWS Direct Connect) would require additional setup and cost, and still limited by the 1 Gbps link.

1114
MCQmedium

A company uses Amazon DynamoDB with global tables in three AWS Regions. The data engineer needs to ensure that writes to the table in us-east-1 are replicated to other regions with minimal latency. Which DynamoDB feature should be used?

A.DynamoDB Global Tables
B.DynamoDB Time to Live (TTL)
C.DynamoDB Streams
D.DynamoDB Accelerator (DAX)
AnswerA

DynamoDB global tables use multi-Region, active-active replication with a last-writer-wins conflict resolution, propagating writes from us-east-1 to the other Regions typically within a second. This directly satisfies the minimal-latency replication constraint, since replication is handled natively by DynamoDB rather than by custom streaming pipelines.

Why this answer

DynamoDB Global Tables is the correct feature because it provides multi-region, multi-master replication, automatically replicating writes from us-east-1 to other regions with sub-second latency. This is achieved through DynamoDB Streams and a last-writer-wins conflict resolution mechanism, ensuring data consistency across regions without requiring custom replication logic.

Exam trap

The trap here is that candidates may confuse DynamoDB Streams (a change capture mechanism) with Global Tables (a managed replication service), not realizing that Streams alone cannot replicate data across regions without additional custom code.

How to eliminate wrong answers

Option B is wrong because DynamoDB Time to Live (TTL) is used to automatically delete expired items based on a timestamp attribute, not for replicating data across regions. Option C is wrong because DynamoDB Streams captures item-level changes in a single table and can trigger AWS Lambda functions, but it does not natively replicate data to other regions; Global Tables uses Streams internally but the feature itself is not a replication solution. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that reduces read latency for a single table, but it does not provide cross-region write replication.

1115
MCQeasy

A data engineer needs to ingest streaming data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real-time and stored in Parquet format for efficient querying. The engineer wants to minimize custom code. Which solution should the engineer use?

A.Use Amazon Kinesis Data Firehose to deliver the stream to S3, and enable record format conversion to Parquet using an AWS Glue table.
B.Use Amazon Kinesis Data Analytics to process the stream and write Parquet to S3 using a Flink application.
C.Use AWS Lambda to read from the Kinesis Data Stream and write Parquet files to S3 using the AWS SDK for pandas (awswrangler).
D.Use AWS Glue streaming ETL job to read from the Kinesis Data Stream and write Parquet to S3.
AnswerA

Kinesis Data Firehose can read directly from a Kinesis Data Stream, buffer records, and deliver them to S3. It supports record format conversion from JSON to Parquet using a schema defined in the AWS Glue Data Catalog. This requires no custom code and provides near real-time delivery, meeting all requirements with minimal effort.

Why this answer

Kinesis Data Firehose provides a fully managed way to deliver streaming data from Kinesis Data Streams to S3. It supports automatic conversion to Parquet using a schema from the AWS Glue Data Catalog, requiring no custom code. This meets the near real-time and Parquet requirements while minimizing development effort, unlike solutions that require writing Lambda functions, Flink applications, or Glue scripts.

Exam trap

The trap here is overlooking Firehose's built-in record format conversion feature and assuming that custom code is needed to convert JSON to Parquet.

1116
MCQhard

A data engineer maintains an AWS Glue ETL job that reads JSON files from Amazon S3, applies transformations using a DynamicFrame, and writes Parquet to another S3 bucket. The job currently reads the entire source prefix on every run, but the source is now partitioned by year/month/day and only new partitions need processing. The engineer wants to process only new partitions without changing the job's transformation logic. Which change should the engineer make?

A.Enable job bookmarks on the Glue job so it tracks previously processed S3 objects and skips them on subsequent runs
B.Create a Glue crawler that catalogs only the new partitions and point the job at the crawler's output table
C.Configure the job to use a smaller number of DPUs so each run scans less data
D.Add a pushdown predicate to the DynamicFrame that filters on the year, month, and day partition columns using hardcoded values
AnswerA

Glue job bookmarks persist state about which S3 objects and partitions were processed in prior runs, so subsequent runs read only new data. This requires no change to the transformation logic and directly addresses incremental processing of newly arriving partitions.

Why this answer

Job bookmarks are the native Glue mechanism for incremental processing. When enabled, Glue tracks the state of previously processed S3 objects and partitions, so each run reads only new data. This meets the requirement to process only new partitions without rewriting the transformation logic or manually specifying dates.

Exam trap

The trap here is confusing partition discovery (crawlers and catalog updates) with partition filtering at read time, when only job bookmarks give automatic incremental reads.

1117
MCQeasy

A data engineer must load data from an Amazon DynamoDB table into an Amazon S3 data lake nightly. The table is approximately 800 GB and the nightly window is tight. The engineer wants a fully managed, serverless option that exports the table to S3 without consuming DynamoDB read capacity or writing custom code. Which solution meets these requirements?

A.Run an AWS Database Migration Service task with DynamoDB as the source and Amazon S3 as the target on a nightly schedule.
B.Use the native Amazon DynamoDB export to Amazon S3 feature, which exports a point-in-time snapshot without consuming read capacity.
C.Use AWS Glue with a DynamoDB connection and the dynamodb.read.throughput.percent option set to 100.
D.Enable DynamoDB Streams, attach an AWS Lambda function, and write each stream record to Amazon S3 as it arrives.
AnswerB

DynamoDB's native export to Amazon S3 produces a full point-in-time snapshot in DynamoDB JSON or Amazon Ion format directly to an S3 bucket. It is serverless, requires no code, and does not consume provisioned or on-demand read capacity, making it ideal for a large nightly export within a tight window.

Why this answer

Amazon DynamoDB's native export to Amazon S3 creates a full point-in-time snapshot of the table directly in an S3 bucket without consuming read capacity and without requiring custom code or replication infrastructure. It is serverless and scales to large tables, which fits the tight nightly window and the no-capacity-impact constraint.

Exam trap

The trap here is assuming DynamoDB Streams can produce a full historical export, when streams only capture changes that occur after the stream is enabled.

1118
MCQhard

A company has a DynamoDB table with a partition key of 'user_id' and a sort key of 'timestamp'. They need to query all items for a user within a date range. Which query operation should be used?

A.BatchGetItem with multiple keys
B.Query with KeyConditionExpression on partition key and sort key
C.GetItem with both partition and sort key
D.Scan with FilterExpression
AnswerB

A Query operation targets a single partition key value and can filter the sort key with a range condition, so KeyConditionExpression on both keys returns exactly the user's items within the date range. Scan cannot filter on the sort key efficiently and reads the whole table.

Why this answer

The Query operation in DynamoDB is designed to retrieve items based on a specific partition key and an optional sort key condition. Since the table has a partition key of 'user_id' and a sort key of 'timestamp', using Query with a KeyConditionExpression that filters on the partition key (user_id) and a range condition on the sort key (timestamp) is the most efficient and correct approach to get all items for a user within a date range.

Exam trap

The trap here is that candidates often confuse BatchGetItem with Query, thinking BatchGetItem can handle range queries, but BatchGetItem only retrieves items by exact primary key values and cannot filter by sort key conditions.

How to eliminate wrong answers

Option A is wrong because BatchGetItem retrieves items by their primary key (partition key and sort key) but does not support range-based filtering on the sort key; it only fetches specific items by exact key values, not a range of timestamps. Option C is wrong because GetItem retrieves a single item by its full primary key (both partition key and sort key), so it cannot return multiple items or filter by a date range. Option D is wrong because Scan reads the entire table and then applies a FilterExpression, which is inefficient and costly for large tables, and it should be avoided when a more targeted Query operation can be used.

1119
MCQmedium

A data engineer is building an AWS Glue ETL job that reads a large Amazon S3 dataset of nested JSON files and must flatten the nested arrays into separate rows for downstream analytics. The engineer needs the most efficient, code-free way to apply this transformation within the Glue job. Which approach should the engineer use?

A.Configure the Glue crawler to automatically flatten nested JSON during cataloging.
B.Use the AWS Glue Studio visual editor and add a Flatten transformation to the job.
C.Write a custom Python shell script in Glue to recursively flatten the JSON.
D.Use an Apache Spark filter transformation to explode the nested arrays.
AnswerB

The Flatten transformation in AWS Glue Studio un-nests nested structures such as arrays and structs into separate rows without writing custom code, which matches the requirement for a code-free, efficient transformation of nested JSON. It is designed exactly for this scenario and integrates with the visual job editor.

Why this answer

The Flatten transformation in AWS Glue Studio is purpose-built to un-nest arrays and structs into rows without custom code, directly satisfying the need for an efficient, code-free flattening step in a Glue ETL job on nested JSON.

Exam trap

The trap here is assuming a crawler or filter can reshape nested data, when only a dedicated Flatten transformation un-nests arrays into rows.

1120
MCQhard

A data engineering team is troubleshooting a slow AWS Glue ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3 in Parquet format. The job processes 50 GB of data. Which action would most effectively improve job performance?

A.Use S3 Select to push down filters
B.Reduce the batch size in the DynamoDB connector
C.Increase the number of DPUs
D.Change output to JSON format to reduce overhead
AnswerC

Adding DPUs scales the number of Apache Spark executors, giving more parallel workers to read the 50 GB DynamoDB table and write Parquet to S3. This directly addresses the compute-bound bottleneck, unlike tuning partition counts or connection settings.

Why this answer

Increasing the number of DPUs (Data Processing Units) for the AWS Glue job directly allocates more distributed computing resources (CPU, memory, and network bandwidth) to parallelize the read from DynamoDB and the write to S3. Since the job processes 50 GB of data, the bottleneck is likely the throughput of the Glue Spark cluster, and adding DPUs increases parallelism, reducing overall execution time.

Exam trap

The trap here is that candidates often confuse S3 Select as a universal filter mechanism or assume that reducing batch size always improves performance, when in fact it increases API overhead and latency in distributed systems like Glue.

How to eliminate wrong answers

Option A is wrong because S3 Select is used to filter data within S3 objects (e.g., CSV or JSON files) and cannot be applied to a DynamoDB source; the filter pushdown must happen at the DynamoDB API level using expressions, not S3 Select. Option B is wrong because reducing the batch size in the DynamoDB connector would decrease the number of items read per request, increasing the number of API calls and likely worsening performance due to higher latency and throttling risk. Option D is wrong because changing output to JSON format would increase file size and write overhead compared to Parquet (which is columnar and compressed), thus degrading performance, not improving it.

1121
MCQhard

A company uses AWS Lake Formation to manage fine-grained access to a data lake in Amazon S3. A data analyst needs to query a table in the AWS Glue Data Catalog that contains columns with sensitive data. The analyst must be able to see only non-sensitive columns and only rows where the region column equals 'US'. The analyst uses Amazon Athena for queries. Which Lake Formation permission model should the data engineer implement?

A.Grant the analyst SELECT permission on the table and use AWS Glue job bookmarks to filter sensitive data.
B.Grant the analyst SELECT permission on the table and create a data filter that includes only non-sensitive columns and a row filter for region = 'US'.
C.Create a view in Athena that selects only non-sensitive columns and filters rows for region = 'US', and grant the analyst access to the view.
D.Grant the analyst SELECT permission on the table and use an IAM policy to deny access to sensitive columns.
AnswerB

Lake Formation data filters allow column-level and row-level security. Granting SELECT on the table with a data filter that specifies the allowed columns and a row filter expression restricts the analyst to only the permitted columns and rows. This meets the requirement precisely without granting broader access.

Why this answer

Lake Formation data filters provide column-level and row-level security by allowing you to specify which columns and rows a principal can access. By granting SELECT with a data filter that includes only non-sensitive columns and a row filter for region = 'US', the analyst is restricted appropriately. IAM policies, Athena views, and Glue job bookmarks do not enforce the required fine-grained access control.

Exam trap

The trap here is assuming that IAM policies or Athena views can enforce column-level and row-level security, when Lake Formation data filters are the correct mechanism for fine-grained access control in a data lake.

1122
MCQeasy

A company wants to ingest streaming data from IoT devices into Amazon S3 using Amazon Kinesis Data Firehose. The data must be transformed from JSON to Parquet format before landing in S3. What is the SIMPLEST way to achieve this?

A.Configure Kinesis Data Firehose with a built-in Parquet converter.
B.Use an AWS Lambda function as a data transformation in Kinesis Data Firehose to convert JSON to Parquet.
C.Use Kinesis Data Firehose to deliver data directly to S3 in JSON format and run a nightly Glue job to convert to Parquet.
D.Use Kinesis Data Analytics to convert the data to Parquet before sending to Firehose.
AnswerA

Kinesis Data Firehose offers a native record format conversion feature that transforms incoming JSON to Parquet using a schema from the AWS Glue Data Catalog, with no custom processing required. This directly satisfies the stem's requirement to convert JSON to Parquet before landing in S3, and remains the simplest option since Firehose handles the transformation itself.

Why this answer

Amazon Kinesis Data Firehose has a built-in Parquet conversion feature that uses an AWS Glue schema to convert incoming JSON data to Parquet format. This is the simplest approach because it requires no custom code or additional services; you only need to provide a schema and enable the conversion in the Firehose delivery stream configuration. Option B (using Lambda) is more complex, as it requires writing and maintaining a custom transformation function.

Exam trap

Candidates often overlook the native Parquet conversion capability in Kinesis Data Firehose and assume a Lambda function is required. The built-in conversion using an AWS Glue schema is actually simpler and supported directly.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Firehose does not have a built-in Parquet converter; it can deliver data in Parquet format only if the input data is already in a format that can be converted (e.g., via a Lambda transformation or by using a schema from AWS Glue), but there is no native 'Parquet converter' toggle. Option C is wrong because it introduces unnecessary complexity and latency by storing JSON in S3 first and then running a nightly Glue job, which is not the simplest approach and does not meet the requirement for real-time transformation before landing. Option D is wrong because Kinesis Data Analytics is designed for real-time analytics and stream processing, not for format conversion; it would add unnecessary overhead and complexity compared to using Firehose's built-in Lambda transformation.

1123
MCQmedium

A data engineer is configuring an S3 bucket policy to allow cross-account access for a partner account to read objects. The bucket is encrypted with SSE-KMS using a customer-managed key. What additional configuration is needed to allow the partner account to decrypt the objects?

A.Add a bucket policy that grants the partner account s3:GetObject
B.Create a VPC endpoint for S3 and add it to the bucket policy
C.Update the KMS key policy to grant the partner account kms:Decrypt permission
D.Add a bucket policy that grants s3:GetObject and s3:GetEncryptionConfiguration
AnswerC

Granting the partner account kms:Decrypt in the KMS key policy satisfies the stem's requirement that cross-account readers decrypt SSE-KMS objects. S3 bucket policies alone cannot authorise key usage; the customer-managed key's policy must separately permit the partner principal, since KMS enforces its own authorisation independently of S3.

Why this answer

For cross-account access with SSE-KMS, the KMS key policy must grant the partner account access to use the key. The bucket policy alone is insufficient. The partner account does not need VPC endpoints, and the bucket policy for decryption is not needed.

The partner account does not need access to the S3 bucket's encryption configuration.

1124
MCQmedium

A data engineer is using AWS Glue Studio to build an ETL job that reads semi-structured JSON from Amazon S3. The source files contain nested arrays and inconsistent keys, and the engineer wants the job to automatically infer the schema at runtime without a Data Catalog table. Which transform or configuration should the engineer use to read the data most reliably?

A.Create a DynamicFrame with the from_options method and set the format to "json" with the recurse parameter enabled.
B.Use a Spark DataFrame read with spark.read.schema() and a hard-coded StructType that matches all nested fields.
C.Configure the job to use the ApplyMapping transform with a manually defined mapping for every nested attribute before reading.
D.Run an AWS Glue crawler against the S3 prefix and then use the catalog table in the job with a from_catalog read.
AnswerA

DynamicFrame.from_options with format "json" and recurse=true handles nested and inconsistent JSON by flattening and inferring the schema dynamically, which suits semi-structured data without a catalog table. It is the standard programmatic approach in AWS Glue for reading JSON from S3 while accommodating schema variability at runtime.

Why this answer

Reading semi-structured JSON from S3 without a Data Catalog entry is best done with DynamicFrame.from_options, specifying the JSON format and enabling recursion so nested arrays and inconsistent keys are handled during inference. This avoids the rigidity of hard-coded schemas and the catalog dependency of crawler-based reads.

Exam trap

The trap here is assuming a Glue crawler or a hard-coded Spark schema is required to read JSON into Glue, when DynamicFrame.from_options with recurse handles runtime inference directly.

1125
MCQmedium

A company uses Amazon S3 to store sensitive data. The data engineer needs to ensure that all data in transit between the S3 bucket and clients is encrypted. Which configuration should the engineer implement?

A.Use Amazon CloudFront to serve the content and enable SSL.
B.Enable default encryption on the S3 bucket using SSE-S3.
C.Create an S3 bucket policy that denies requests where SecureTransport is false.
D.Use SSE-C to encrypt the data with a customer-provided key.
AnswerC

The aws:SecureTransport condition key evaluates whether the request arrived over TLS; denying when it is false blocks any plain HTTP access to objects. This enforces encryption in transit between clients and the bucket, which is exactly the stem's stated requirement.

Why this answer

An S3 bucket policy that denies requests where SecureTransport is false enforces HTTPS for all access, encrypting data in transit. Option A is wrong because CloudFront with SSL can enforce HTTPS but is not a direct S3 configuration and may incur additional costs. Option B is wrong because SSE-S3 only encrypts data at rest, not in transit.

Option D is wrong because SSE-C also encrypts data at rest with a customer-provided key, not in transit.

Page 14

Page 15 of 18

Page 16