Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 1276–1321

1321 questions total · 18pages · All types, answers revealed

Page 17

Page 18 of 18

1276
MCQhard

A data engineer is troubleshooting an issue where an Amazon Redshift query returns an error: 'ERROR: permission denied for relation table_name'. The user has been granted SELECT on the table. What is the most likely cause?

A.The user's session has timed out.
B.The user does not have CONNECT permission on the database.
C.The table is in a different schema than expected.
D.The user does not have USAGE permission on the schema.
AnswerD

Redshift requires USAGE on the containing schema before SELECT on a table is honoured. Without schema USAGE, the query fails with permission denied even though SELECT was granted, making the missing schema-level privilege the most likely cause.

Why this answer

In Amazon Redshift, to access a table, a user must have USAGE permission on the schema containing the table, in addition to SELECT or other table-level permissions. Without USAGE on the schema, the user receives a 'permission denied for relation' error even if SELECT is granted. Option D is correct.

Option A (session timeout) would cause a different error or disconnection. Option B (no CONNECT permission) would prevent connecting to the database. Option C (wrong schema) would result in a 'schema not found' error, not a permission denied error.

1277
MCQmedium

A data engineer is using AWS Lake Formation to manage access to a data lake in Amazon S3. The engineer needs to grant a specific IAM role access to only the columns containing non-sensitive data in a table stored in the AWS Glue Data Catalog. The role should not have access to sensitive columns. What should the engineer do?

A.Grant the IAM role SELECT permission on the table, and then apply an IAM policy that denies access to the columns with sensitive data.
B.Use AWS Glue Data Catalog resource policies to deny access to the sensitive columns for the IAM role.
C.Create a Lake Formation data filter that excludes the sensitive columns, and grant the IAM role SELECT permission on the table with the data filter applied.
D.Create a view in Amazon Athena that selects only the non-sensitive columns, and grant the IAM role access to the view instead of the table.
AnswerC

Lake Formation data filters allow column-level and row-level access control. By creating a data filter that excludes sensitive columns and granting SELECT with that filter, the IAM role can access only the non-sensitive columns. This is the intended way to implement fine-grained access control in Lake Formation and meets the requirement precisely.

Why this answer

AWS Lake Formation provides column-level security through data filters. By creating a data filter that excludes sensitive columns and granting SELECT with that filter, the engineer ensures the IAM role can only access the permitted columns. This is the native and most secure method for fine-grained access control in Lake Formation.

Exam trap

The trap here is assuming that IAM policies or Glue Data Catalog resource policies can enforce column-level permissions, when in fact only Lake Formation data filters provide that capability.

1278
MCQeasy

A logistics company stores shipment tracking events in an Amazon DynamoDB table. The table uses a partition key of shipment_id and a sort key of event_timestamp. Analysts frequently run queries that filter by shipment_id and a range of event_timestamp values. The data engineer must ensure these queries are efficient and consume minimal read capacity. What should the data engineer do?

A.Use the base table with a Query operation specifying shipment_id as the partition key and a condition on event_timestamp.
B.Create a global secondary index on event_timestamp alone.
C.Run a Scan operation with a FilterExpression on shipment_id and event_timestamp.
D.Enable DynamoDB Streams on the table and query the stream for matching events.
AnswerA

The base table's composite primary key of shipment_id and event_timestamp is exactly the shape needed. A Query with shipment_id as the partition key and a range condition on event_timestamp reads only the matching item collection, avoids full table scans, and consumes read capacity proportional to the items returned.

Why this answer

The composite primary key already matches the access pattern, so a Query with shipment_id as the partition key and a condition on event_timestamp reads only the relevant item collection. This is the most efficient and lowest-cost approach and requires no additional index or infrastructure.

Exam trap

The trap here is reaching for a secondary index or Streams when the existing primary key already supports the query pattern efficiently.

1279
MCQeasy

A data engineer needs to store time-series data from IoT devices. The data is write-heavy and requires low-latency queries by device ID and timestamp. The data volume is expected to grow to terabytes. Which AWS database service is most suitable?

A.Amazon RDS for MySQL
B.Amazon ElastiCache for Redis
C.Amazon DynamoDB
D.Amazon Timestream
AnswerD

Amazon Timestream is purpose-built for time-series workloads, using a memory store for recent data and a magnetic store for historical data. This satisfies the stem's write-heavy, low-latency-by-device-and-timestamp, terabyte-scale requirements, which general-purpose relational or key-value services cannot match as efficiently.

Why this answer

Amazon Timestream is purpose-built for time-series data, offering automatic tiered storage (in-memory for recent data and magnetic for historical) to handle write-heavy IoT workloads at scale. It supports low-latency queries by device ID and timestamp via its SQL-compatible query engine, making it the most suitable choice for terabytes of time-series data.

Exam trap

The trap here is that candidates often choose DynamoDB (Option C) because of its high write throughput and low-latency queries, but they overlook the lack of native time-series optimizations, leading to complex manual partitioning and TTL management that Timestream handles automatically.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for MySQL is a relational database optimized for OLTP workloads with structured queries, not for the high-volume, write-heavy, time-series pattern that requires automatic data retention policies and time-based partitioning. Option B is wrong because Amazon ElastiCache for Redis is an in-memory cache designed for sub-millisecond read/write performance on hot data, but it cannot cost-effectively store terabytes of data and lacks native time-series query optimizations like downsampling and interpolation. Option C is wrong because Amazon DynamoDB is a key-value and document database that can handle high write throughput, but it does not have built-in time-series functions (e.g., time-based aggregation, retention policies) and requires manual partitioning and TTL management to handle time-series data efficiently at terabyte scale.

1280
MCQhard

A company has an Amazon RDS for MySQL database that is experiencing performance issues due to a large number of read requests. The application is read-heavy and can tolerate eventually consistent reads. Which action will reduce the load on the primary database with the least operational overhead?

A.Create a read replica in the same region
B.Use Amazon ElastiCache for caching
C.Enable Multi-AZ deployment
D.Increase the instance size of the primary DB
AnswerA

A read replica asynchronously replicates from the primary and serves read-only traffic, offloading read requests. Because the application tolerates eventually consistent reads, this fits, and creating a replica requires minimal operational effort compared with sharding or caching layers.

Why this answer

Creating a read replica in the same region offloads read traffic from the primary RDS instance to a read-only copy, which directly addresses the read-heavy workload. Since the application can tolerate eventually consistent reads, the slight replication lag is acceptable, and this solution requires minimal operational overhead—just a few clicks in the AWS console or a single API call—without any application code changes.

Exam trap

The trap here is that candidates often confuse Multi-AZ (which only provides failover, not read scaling) with read replicas, or assume that ElastiCache is always the best caching solution without considering the operational overhead of code changes and cache management.

How to eliminate wrong answers

Option B is wrong because using Amazon ElastiCache introduces additional infrastructure to manage (e.g., cache invalidation, cluster configuration) and requires application code changes to implement caching logic, increasing operational overhead compared to a simple read replica. Option C is wrong because enabling Multi-AZ deployment provides high availability and automatic failover but does not offload read traffic; the standby instance is not used for reads, so it does not reduce load on the primary. Option D is wrong because increasing the instance size of the primary DB only scales vertically, which can be costly and still leaves all read traffic hitting a single instance, failing to distribute the load and not leveraging the read-heavy, eventually consistent tolerance.

1281
Multi-Selectmedium

A data engineer needs to ensure that data in an Amazon S3 bucket is not publicly accessible. Which TWO measures should the engineer implement? (Choose TWO.)

Select 2 answers
A.Attach a bucket policy that denies access to 'Principal': '*' unless specific conditions are met.
B.Create a lifecycle policy to delete objects after 30 days.
C.Enable S3 Block Public Access settings on the bucket.
D.Enable S3 Versioning on the bucket.
E.Enable default encryption on the bucket.
AnswersA, C

A bucket policy denying Principal '*' blocks anonymous and cross-account public access, satisfying the requirement that objects remain non-public. Adding conditions restricts even authenticated principals, closing the common misconfiguration where ACLs or Block Public Access alone leave policy-based exposure.

Why this answer

Option A is correct because an S3 bucket policy with an explicit Deny for 'Principal': '*' (optionally scoped by conditions such as aws:SourceIp or aws:SourceVpce) overrides any Allow and prevents anonymous or public access to the bucket's objects. Option C is correct because S3 Block Public Access provides account-level and bucket-level controls (BlockPublicAcls, IgnorePublicAcls, BlockPublicPolicy, RestrictPublicBuckets) that reject public ACLs and public bucket policies, which is the primary AWS safeguard against accidental public exposure. Option B is incorrect because a lifecycle policy only transitions or expires objects and has no effect on access permissions.

Option D is incorrect because S3 Versioning only retains multiple object versions for recovery and does not restrict who can read the data. Option E is incorrect because default encryption protects data at rest but does not prevent public read access to the bucket or its objects.

1282
Multi-Selectmedium

A data engineer must give an AWS Glue ETL job access to an S3 bucket that is encrypted with SSE-KMS using a customer managed key. The Glue job runs under an IAM role. The security team wants the least-privilege permissions required for the job to read and write objects in that bucket. Which TWO actions must be included in the IAM role's policy? (Choose two.)

Select 2 answers
A.kms:ListKeys on the customer managed key.
B.kms:Decrypt on the customer managed key.
C.kms:GenerateDataKey on the customer managed key.
D.kms:ScheduleKeyDeletion on the customer managed key.
E.kms:CreateGrant on the customer managed key.
AnswersB, C

When S3 objects are encrypted with SSE-KMS, reading an object requires the caller to have kms:Decrypt on the key that protects the object. The Glue job role must include this action or S3 returns AccessDenied during the read. Without it, the job cannot decrypt the data even if it has s3:GetObject.

Why this answer

Reading SSE-KMS encrypted objects requires kms:Decrypt on the key, and writing them requires kms:GenerateDataKey so S3 can obtain a fresh data key per object. These two KMS actions, combined with the appropriate s3:GetObject and s3:PutObject permissions, give the Glue job the minimum cryptographic access it needs without granting administrative key management capabilities.

Exam trap

The trap here is forgetting that SSE-KMS adds KMS permissions on top of S3 permissions, so an S3-only policy will still fail with AccessDenied.

1283
MCQeasy

A data engineer is designing a data pipeline that ingests data from an on-premises database into Amazon S3 using AWS Database Migration Service (DMS). The data must be encrypted at rest in S3 using SSE-S3. The engineer also needs to track changes to the source database in real time. Which DMS configuration should the engineer use?

A.Use DMS with a snapshot of the source database.
B.Use DMS with ongoing replication (change data capture) enabled.
C.Use DMS with a full load task only.
D.Use DMS with a full load task and then stream to Amazon Kinesis.
AnswerB

Enabling change data capture with ongoing replication lets DMS read source transaction logs and apply continuous changes to S3, satisfying the real-time tracking requirement. Full load alone captures only a point-in-time snapshot, so CDC is the mechanism that meets the stem's constraint.

Why this answer

DMS with ongoing replication (change data capture) enables real-time tracking of changes from the source database. Option A is incorrect because using a snapshot only captures data at a point in time, not real-time changes. Option C is incorrect because a full load task only loads existing data without capturing ongoing changes.

Option D is incorrect because streaming to Amazon Kinesis is unnecessary; DMS CDC can directly replicate changes to S3. Encryption at rest in S3 with SSE-S3 is automatically supported by DMS when writing to S3.

1284
MCQmedium

A company is running an Amazon EMR cluster with Spark for data processing. The data engineer wants to automatically scale the core and task nodes based on the YARN memory and CPU utilization. Which scaling metric should the engineer use for the EMR managed scaling policy?

A.YARNMemoryAvailablePercentage
B.CPUUtilization
C.DiskIOPS
D.HDFSUtilization
AnswerA

YARNMemoryAvailablePercentage reflects the memory available to YARN containers, which governs whether Spark executors can be scheduled on core and task nodes. Scaling on this metric adds or removes nodes in step with actual workload demand, matching the stem's YARN memory and CPU utilisation requirement.

Why this answer

EMR managed scaling policies are designed around YARN metrics, and YARNMemoryAvailablePercentage is the primary metric used to determine when to add or remove core and task nodes. When available YARN memory drops below a threshold, EMR scales out; when it rises above a threshold, EMR scales in. This directly reflects the resource pressure that Spark executors place on the cluster.

Exam trap

The trap is picking a familiar CloudWatch metric like CPUUtilization or DiskIOPS instead of the YARN-centric metric that EMR managed scaling actually uses; candidates who have not read the EMR managed scaling documentation often fall for this.

How to eliminate wrong answers

Option B is wrong because CPUUtilization is a CloudWatch metric that does not capture YARN container allocation and is not the metric EMR managed scaling uses to decide scale actions. Option C is wrong because DiskIOPS measures storage throughput and is not a scaling trigger for EMR managed scaling, which is memory-driven. Option D is wrong because HDFSUtilization reflects HDFS storage capacity, relevant for HDFS-heavy workloads but not the metric EMR managed scaling uses to scale core and task nodes based on YARN memory and CPU.

1285
Multi-Selecteasy

A data engineer is monitoring an Amazon Kinesis Data Stream used to ingest clickstream data. The engineer notices that the stream's 'WriteProvisionedThroughputExceeded' metric is frequently above zero. Which TWO actions could help mitigate this issue? (Choose TWO.)

Select 2 answers
A.Increase the number of shards in the stream.
B.Reduce the data retention period to free up capacity.
C.Decrease the number of shards to reduce overhead.
D.Implement a random prefix for the partition key to distribute data evenly.
E.Enable enhanced fan-out on the stream.
AnswersA, D

Adding shards raises the stream's total ingest capacity, since each shard provides a fixed write throughput ceiling (1 MB/s, 1,000 records/s). The stem's persistent WriteProvisionedThroughputExceeded metric indicates producers are exceeding provisioned capacity, so scaling shard count directly relieves that constraint.

Why this answer

Option A is correct because WriteProvisionedThroughputExceeded indicates that producers are exceeding the stream's ingest capacity, and each shard supports a fixed write throughput of 1 MB/s or 1,000 records/s, so adding shards (resizing/shard splitting) increases total write capacity. Option D is correct because a hot shard caused by an uneven partition key (for example, a constant or low-cardinality key) can throttle writes even when aggregate capacity is sufficient; adding a random prefix to the partition key spreads records across more shards, balancing the load. Option B is incorrect because the retention period only controls how long data is stored and has no effect on write throughput capacity.

Option C is incorrect because decreasing shards reduces total write capacity and would worsen throttling. Option E is incorrect because enhanced fan-out increases read throughput for consumers (dedicated 2 MB/s per consumer per shard) and does not address write-side throttling.

Exam trap

DEA-C01 often tests the misconception that read-side features like enhanced fan-out or retention settings can fix producer-side throttling — candidates must separate write capacity (shards, partition key distribution) from read capacity (fan-out, consumers).

1286
MCQhard

A data engineer is managing an Amazon S3 data lake that contains millions of small files. The engineer needs to optimize query performance in Amazon Athena and reduce costs. The data is stored in Parquet format and is partitioned by date. Which action should the engineer take to improve performance and reduce costs?

A.Use AWS Glue ETL to compact small files into larger files.
B.Convert the data to CSV format to enable faster scans.
C.Enable S3 Versioning on the bucket to improve read performance.
D.Increase the number of partitions to reduce the amount of data scanned per query.
AnswerA

Compacting small files into larger files reduces the number of S3 GET requests and metadata overhead, improving Athena query performance. It also reduces the amount of data scanned if the compaction results in better compression and columnar storage. This is a common optimization for data lakes with many small files.

Why this answer

Compacting small files into larger files using AWS Glue ETL reduces the number of files and S3 requests, improving Athena query performance and reducing costs. Other options either do not address the small file issue or would degrade performance. This is a best practice for optimizing data lakes.

Exam trap

The trap here is thinking that adding more partitions or enabling S3 Versioning will help, when the core issue is the number of small files, which requires compaction.

1287
Multi-Selecthard

A company is migrating a legacy data warehouse to Amazon Redshift. They need to choose a distribution style to minimize data movement during joins. Which THREE factors should they consider?

Select 3 answers
A.The size of the table (number of rows).
B.The join frequency with other tables on specific columns.
C.The number of columns in the table.
D.Whether the table is a fact or dimension table.
E.The data type of the distribution key column.
AnswersA, B, D

Table size determines whether a table qualifies for ALL distribution, which broadcasts a small table to every node and eliminates join data movement entirely. For large tables, row count instead guides choosing DISTKEY on the frequent join column, since ALL replication becomes impractical. Both paths directly minimise the internode traffic the stem requires.

Why this answer

Option A is correct because table size (number of rows) is a primary driver of distribution style choice: small tables are typically assigned ALL distribution so they can be broadcast to every node, eliminating data movement during joins, while large tables need KEY or EVEN distribution. Option B is correct because the whole purpose of choosing a distribution style is to co-locate matching rows on the same slice; if a table is frequently joined on a specific column, distributing both tables on that join key keeps the join local and avoids network redistribution. Option D is correct because fact and dimension tables play different roles in a star schema: large fact tables are usually distributed on their most common join key (or EVEN), while smaller dimension tables are often set to ALL so they are replicated to every node and can be joined without movement.

Option C is not a factor because the number of columns does not affect how rows are distributed across slices; it only affects storage width and I/O, not join data movement. Option E is not a factor because Redshift supports distribution keys of various data types, and the data type itself does not determine whether a join requires redistribution — what matters is whether the joined columns match and are used as distribution keys.

Exam trap

The trap here is that candidates may overthink irrelevant table properties like column count or data types, while the core considerations for minimizing data movement are table size, join frequency, and table role (fact vs. dimension).

1288
MCQmedium

A data engineer is building a data pipeline that ingests data from Amazon S3 into Amazon Redshift. The data is in CSV format and includes a timestamp column. The pipeline should load only new data incrementally. Which approach is most efficient?

A.Use the COPY command to load the entire bucket and rely on Redshift to deduplicate
B.Use the COPY command with a manifest file that lists only the new S3 objects
C.Use Amazon Redshift Spectrum to query the S3 data directly without loading
D.Use INSERT statements within a loop to load each new file
AnswerB

A manifest listing only new S3 objects lets COPY load precisely those files, avoiding rescanning or reprocessing previously ingested data. This satisfies the incremental-load constraint efficiently, since Redshift reads just the specified object set rather than filtering the whole bucket.

Why this answer

Using the COPY command with a manifest file allows you to explicitly list only the new S3 objects to be loaded, enabling incremental loading without scanning or loading the entire bucket. This approach is efficient as it avoids the overhead of deduplication or full-bucket scans, and it leverages Redshift's native high-speed parallel ingestion from S3.

Exam trap

The trap here is that candidates may think Redshift Spectrum is a valid alternative for loading data, but Spectrum is designed for external querying, not for persistent loading into Redshift tables, which is the explicit requirement in the question.

How to eliminate wrong answers

Option A is wrong because loading the entire bucket and relying on Redshift to deduplicate is inefficient; Redshift does not have built-in deduplication logic for COPY commands, and loading all data repeatedly would waste storage and compute resources. Option C is wrong because Redshift Spectrum queries data directly in S3 without loading it into Redshift tables, which does not meet the requirement of loading data into Redshift for persistent storage and incremental processing. Option D is wrong because using INSERT statements within a loop to load each new file is far less efficient than the COPY command, as it lacks parallelization and high-throughput optimization, leading to poor performance for large datasets.

1289
MCQeasy

A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is consumed by a custom consumer application that writes to Amazon S3 every 5 minutes. The consumer is falling behind and processing lag is increasing. Which action is MOST effective to reduce the lag?

A.Switch to Amazon Kinesis Data Firehose to deliver data directly to S3
B.Increase the batch size of records written to S3
C.Increase the number of shards in the Kinesis stream
D.Reduce the retention period of the stream
AnswerC

A Kinesis stream's throughput ceiling is set by shard count: each shard provides 1 MB/s or 1,000 records/s ingest and 2 MB/s egress. Adding shards raises parallel capacity so the consumer can drain the backlog faster, directly reducing processing lag.

Why this answer

The consumer is falling behind because the stream's throughput capacity is insufficient for the incoming data volume. Increasing the number of shards in the Kinesis stream directly increases the total read capacity (each shard provides 2 MB/s read throughput and 5 transactions/second), allowing the consumer to process more data in parallel and reduce lag.

Exam trap

The trap here is that candidates often confuse throughput scaling with batch size or delivery destination changes, but the only way to increase read throughput from a Kinesis stream is to increase the number of shards or use enhanced fan-out.

How to eliminate wrong answers

Option A is wrong because switching to Kinesis Data Firehose does not change the underlying stream's throughput; Firehose is a delivery service that still reads from the same shards, so it would not resolve the consumer's processing lag. Option B is wrong because increasing the batch size written to S3 only affects the write operation to S3, not the consumer's ability to read from the stream faster; the bottleneck is the consumer's read throughput, not the S3 write batch size. Option D is wrong because reducing the retention period (default 24 hours to 1 hour) does not increase read throughput; it only causes data to expire sooner, which could lead to data loss but does not help the consumer catch up.

1290
MCQmedium

A data engineer is configuring an AWS Glue ETL job that reads semi-structured JSON event logs from Amazon S3 and must flatten nested arrays into relational columns before writing to Amazon Redshift. The job must run reliably without writing custom serialization code. Which approach should the data engineer take?

A.Use the AWS Glue DynamicFrame Relationalize transform to convert nested structures into separate relational tables before loading into Redshift.
B.Enable the AWS Glue job bookmark and rely on the crawler's inferred schema to automatically normalize nested JSON during the write to Redshift.
C.Apply the AWS Glue ApplyMapping transform to rename nested fields and then write the DynamicFrame directly to Redshift.
D.Convert each JSON file to CSV using an AWS Lambda function triggered by S3 event notifications before the Glue job runs.
AnswerA

Relationalize is purpose-built for flattening nested JSON into a set of related tables, handling arrays and structs automatically without custom code. It produces a root table plus child tables keyed by generated join identifiers, which can then be written to Redshift. This matches the requirement to avoid custom serialization.

Why this answer

The Relationalize transform is the AWS Glue built-in that flattens nested JSON into a root table and child tables linked by join keys, precisely the normalization needed before loading into a relational target like Redshift. It removes the need to hand-code serialization, keeping the pipeline maintainable and aligned with serverless Glue patterns.

Exam trap

The trap here is assuming that any Glue transform can flatten nested JSON, when only Relationalize is designed to explode arrays and structs into related relational tables.

1291
MCQeasy

A company uses Amazon Redshift for its data warehouse. The data engineering team loads data daily from Amazon S3 using COPY commands. Recently, the load performance has degraded because the S3 bucket contains many small files. The team needs to optimize the COPY operation to improve performance. Which approach should they take?

A.Use Redshift Spectrum to query the data directly from S3 without loading.
B.Increase the number of nodes in the Redshift cluster.
C.Use a manifest file that lists only the necessary files, and consolidate small files into larger ones before loading.
D.Enable automatic compression on the Redshift table.
AnswerC

A manifest restricts COPY to the listed files, avoiding the per-file overhead that many small objects cause, while consolidating them into larger files reduces the number of COPY operations. Together these directly address the degraded load performance.

Why this answer

The performance degradation is caused by the overhead of processing many small files during the COPY command. Consolidating small files into larger ones (e.g., 100 MB–1 GB each) reduces the number of S3 GET requests and the metadata overhead on Redshift, directly improving load throughput. Using a manifest file further optimizes by explicitly listing only the required files, avoiding unnecessary S3 list operations.

Exam trap

The trap here is that candidates often confuse scaling the cluster (Option B) with optimizing data ingestion, failing to recognize that the bottleneck is the number of S3 objects, not the cluster's compute capacity.

How to eliminate wrong answers

Option A is wrong because Redshift Spectrum queries data in place from S3 without loading it into Redshift tables, which does not optimize the COPY operation for loading data into the warehouse. Option B is wrong because increasing the number of nodes adds compute and storage capacity but does not address the root cause of many small files; the COPY command still suffers from the same per-file overhead regardless of cluster size. Option D is wrong because automatic compression (via the COPY command with the COMPUPDATE option) optimizes column encoding for storage efficiency, not the file-level I/O performance during the load process.

1292
MCQeasy

A data engineer is troubleshooting a failed AWS Glue Crawler. The crawler logs show 'Insufficient permissions to access S3 bucket'. What should the engineer do to resolve this?

A.Grant the crawler's IAM user access to the bucket
B.Attach a VPC endpoint to the S3 bucket
C.Enable S3 default encryption on the bucket
D.Update the IAM role used by the crawler to include S3 read permissions
AnswerD

The crawler assumes an IAM role to read S3 data; its current role lacks the required s3:GetObject and s3:ListBucket permissions. Attaching an S3 read policy to that role directly resolves the logged authorisation failure, satisfying the crawler's need to access the source bucket.

Why this answer

The AWS Glue Crawler uses an IAM role to access data sources. The error 'Insufficient permissions to access S3 bucket' indicates that the IAM role attached to the crawler lacks the necessary S3 read permissions (e.g., s3:GetObject, s3:ListBucket). Updating the IAM role's policy to include these permissions resolves the issue, as the crawler operates under that role, not under a specific IAM user.

Exam trap

The trap here is that candidates may confuse the crawler's execution context with an IAM user, leading them to choose Option A, but AWS Glue Crawlers always run under an IAM role, not a user.

How to eliminate wrong answers

Option A is wrong because AWS Glue Crawlers do not use an IAM user for execution; they use an IAM role. Granting access to an IAM user would not affect the crawler's permissions. Option B is wrong because a VPC endpoint enables private connectivity between a VPC and S3 but does not grant or modify IAM permissions; the error is about authorization, not network connectivity.

Option C is wrong because enabling S3 default encryption controls server-side encryption settings and does not affect IAM permission policies; the crawler still needs explicit read access regardless of encryption.

1293
MCQeasy

A company has a nightly batch job that processes 100 GB of data from an Amazon S3 bucket and loads it into an Amazon Redshift table. The job currently runs on an Amazon EMR cluster. Which service would reduce operational overhead while providing similar functionality?

A.AWS Database Migration Service
B.AWS Glue
C.Amazon Redshift Spectrum
D.Amazon Athena
AnswerB

AWS Glue provides serverless Spark-based ETL, eliminating cluster provisioning and tuning that EMR requires. It reads directly from S3 and writes to Redshift, delivering comparable transformation capability for the 100 GB nightly batch while removing the operational burden of managing nodes.

Why this answer

AWS Glue is a serverless ETL service that can process 100 GB of data from S3 and load it into Redshift without managing any infrastructure. It provides built-in job scheduling, automatic retries, and a Spark-based engine that handles large-scale data transformations, directly replacing the EMR cluster's functionality while eliminating operational overhead.

Exam trap

The DEA-C01 exam often tests the distinction between query engines (Athena, Redshift Spectrum) and ETL services (Glue), where candidates mistakenly choose Athena or Spectrum because they can read from S3, but they lack the batch processing and data loading capabilities required for this use case.

How to eliminate wrong answers

Option A is wrong because AWS Database Migration Service (DMS) is designed for continuous database replication or one-time migrations between databases, not for batch processing and transforming large datasets from S3 into Redshift. Option C is wrong because Amazon Redshift Spectrum allows querying data directly in S3 without loading it into Redshift, but it does not perform ETL transformations or replace the batch job's processing logic. Option D is wrong because Amazon Athena is an interactive query service for ad-hoc analysis on S3 data, not a managed ETL service for scheduled batch processing and loading into Redshift.

1294
MCQmedium

A data engineer is running an Amazon Athena query that scans a large amount of data in Amazon S3, resulting in high costs. The data is stored in Parquet format in a partitioned table. Which strategy would be MOST effective in reducing the amount of data scanned?

A.Ensure the query includes a WHERE clause that filters on partition columns.
B.Convert the Parquet files to CSV format and apply GZIP compression.
C.Use S3 Intelligent-Tiering storage class to reduce storage costs.
D.Increase the number of partitions by adding more partition columns.
AnswerA

Partition pruning uses the WHERE clause on partition columns to skip entire partitions of the Parquet table, so Athena reads only matching S3 prefixes. This directly reduces bytes scanned, which is the cost driver for the query.

Why this answer

In Athena, partition pruning is the single most effective way to reduce data scanned because Athena only reads the S3 prefixes that match the partition filter. Adding a WHERE clause on partition columns (A) lets the query engine skip entire partitions, dramatically lowering both cost and runtime. This is especially impactful with Parquet, which already supports columnar projection and predicate pushdown.

Exam trap

The trap is assuming that compression or storage-class optimization reduces Athena query cost — candidates forget that Athena bills on bytes scanned, so only partition pruning and columnar formats actually lower the bill.

How to eliminate wrong answers

Option B is wrong because converting Parquet to CSV removes columnar storage and predicate pushdown, which would increase data scanned, not reduce it — GZIP compression reduces storage size but not the bytes Athena must read for a query. Option C is wrong because S3 Intelligent-Tiering only optimizes storage cost, not query cost; Athena charges per byte scanned regardless of storage class. Option D is wrong because adding more partition columns without a matching filter in the query does not reduce scanned data and can actually create the small-files problem, increasing overhead.

1295
MCQhard

A data engineer is designing a DynamoDB table for an application that requires strongly consistent reads and supports a global secondary index (GSI). The engineer needs to ensure that queries on the GSI return the most up-to-date data. Which statement about DynamoDB read consistency is correct?

A.GSI queries can be strongly consistent only if the index is created with a specific attribute for versioning.
B.GSI queries always return strongly consistent data because the index is maintained synchronously.
C.GSI queries are eventually consistent and cannot be made strongly consistent.
D.GSI queries can be strongly consistent if the base table uses on-demand capacity mode.
AnswerC

DynamoDB GSIs are maintained asynchronously from the base table, so queries on a GSI are eventually consistent and may return stale data. Strongly consistent reads are not supported on GSIs. This is a fundamental limitation, so the engineer must accept eventual consistency or query the base table directly for strong consistency.

Why this answer

DynamoDB GSIs are updated asynchronously from the base table, so reads from a GSI are eventually consistent and cannot be made strongly consistent. The engineer must either accept eventual consistency or query the base table directly for strongly consistent reads. This is a key design constraint when using GSIs.

Exam trap

The trap here is assuming that GSI reads can be strongly consistent like base table reads, when GSIs are inherently eventually consistent.

1296
Multi-Selecthard

A company is running a Redshift cluster and wants to improve query performance for a frequently used dashboard. Which THREE approaches are recommended?

Select 3 answers
A.Enable concurrency scaling
B.Apply column compression encoding
C.Define sort keys on columns used in WHERE clauses
D.Add more nodes to the cluster
E.Choose an appropriate distribution key for large tables
AnswersB, C, E

Column compression encoding reduces the bytes read from disk per column, so scans for dashboard queries retrieve less data and consume less I/O bandwidth. This directly accelerates the frequently executed queries by shrinking storage footprint and improving scan efficiency.

Why this answer

Option B is correct because applying column compression encoding reduces the amount of data read from disk, which lowers I/O and speeds up scans for dashboard queries. Option C is correct because defining sort keys on columns used in WHERE clauses enables zone maps to skip irrelevant blocks, dramatically reducing the data scanned for filtered queries. Option E is correct because choosing an appropriate distribution key for large tables colocates related rows on the same node slice, minimizing data movement during joins and aggregations common in dashboards.

Option A is not among the marked answers, and concurrency scaling mainly handles concurrent workload spikes rather than improving single-query performance. Option D is not marked because simply adding nodes increases capacity and cost but does not by itself optimize query execution without proper sort, distribution, and encoding design.

1297
Multi-Selecteasy

Which TWO AWS services can be used to ingest streaming data into Amazon S3? (Choose two.)

Select 2 answers
A.Amazon S3 Transfer Acceleration
B.Amazon Managed Streaming for Apache Kafka (Amazon MSK)
C.Amazon Kinesis Data Firehose
D.Amazon Elastic Block Store (Amazon EBS)
E.AWS Snowball
AnswersB, C

Amazon MSK runs Apache Kafka clusters whose topics hold streaming records; consumers, or Kafka Connect's S3 sink connector, read those topics and write objects into Amazon S3. This satisfies the stem's requirement for a service that ingests streaming data into S3.

Why this answer

Amazon Kinesis Data Firehose is the easiest way to reliably load streaming data into Amazon S3. It can capture, transform, and deliver streaming data to S3 destinations in near real-time with no code required. Amazon MSK (Managed Streaming for Apache Kafka) can also ingest streaming data into S3 by using Kafka Connect with an S3 sink connector, which writes data from Kafka topics directly to S3.

Exam trap

The trap here is that candidates often confuse Amazon S3 Transfer Acceleration (a speed optimization for existing uploads) with a streaming ingestion service, or they mistakenly think EBS or Snowball can handle real-time streaming data when they are designed for persistent block storage and offline bulk transfer, respectively.

1298
Multi-Selecteasy

A data engineer is designing a real-time streaming pipeline to ingest clickstream data from a website into Amazon S3. The data must be transformed before storage. Which TWO AWS services can be used together to build this pipeline? (Choose TWO.)

Select 2 answers
A.Amazon Kinesis Data Firehose
B.AWS Glue
C.Amazon Kinesis Data Streams
D.Amazon S3 Transfer Acceleration
E.AWS Database Migration Service (DMS)
AnswersA, C

Delivers streaming data to S3 with transformation capabilities.

Why this answer

The correct combination is Amazon Kinesis Data Streams to ingest the streaming clickstream data, and Amazon Kinesis Data Firehose to read from the stream, perform transformations (e.g., via Lambda), and deliver the transformed data to Amazon S3. Kinesis Data Streams provides the durable ingestion layer, while Kinesis Data Firehose handles the delivery and transformation without managing servers.

Exam trap

The trap here is that candidates often confuse Kinesis Data Streams (which requires custom consumers and does not natively write to S3) with Kinesis Data Firehose (which directly delivers to S3 and supports built-in transformations), leading them to select only Data Streams or miss the need for a transformation service.

1299
MCQhard

A data engineer is using AWS Glue to read a large dataset from Amazon S3 and write it to Amazon Redshift. The job intermittently fails with 'Communication link failure' errors during the write phase. The dataset is several hundred gigabytes and the Redshift cluster is under heavy query load. Which change is MOST likely to resolve the failures while preserving data integrity?

A.Switch the write operation to use Redshift Spectrum and query the S3 data directly without loading.
B.Write the transformed data to Amazon S3 in Parquet, then use a Redshift COPY command from S3 instead of a direct JDBC write.
C.Increase the Glue job's number of workers and enable auto-scaling.
D.Configure the Glue connection to use a Redshift JDBC URL with a longer socketTimeout and retry the write.
AnswerB

The Glue Redshift connector can stage data in S3 and issue a COPY command, which is far more resilient and efficient than row-by-row JDBC inserts. COPY is atomic per transaction, handles large volumes well, and reduces the number of open connections to Redshift. This approach preserves data integrity and avoids the communication link failures caused by direct writes under load.

Why this answer

Staging data in S3 and using the Redshift COPY command is the recommended pattern for bulk loads. It minimizes open connections, leverages Redshift's parallel load capability, and provides transactional integrity. Increasing workers, extending timeouts, or switching to Spectrum does not resolve the underlying contention or integrity concerns during direct writes.

Exam trap

The trap here is treating the communication failure as a Glue-side timeout or capacity problem, when the real fix is to change the write path to use S3 staging plus Redshift COPY.

1300
MCQmedium

A company uses AWS Glue ETL jobs to transform data from Amazon S3 to Amazon Redshift. The job reads JSON files, applies schema mapping, and writes to a Redshift table. Recently, the job started failing with memory errors. The data volume has increased tenfold. Which approach should a data engineer take to resolve this issue with minimal code changes?

A.Switch from Spark to Python Shell job type.
B.Implement batch processing with smaller file sizes.
C.Increase the number of DPUs allocated to the Glue job.
D.Use Redshift Spectrum to query data directly from S3.
AnswerC

Glue allocates memory per executor, so raising the DPU count adds workers and distributes the tenfold-larger dataset across more memory, resolving the out-of-memory failures. It requires only a job configuration change, meeting the minimal-code-change constraint.

Why this answer

Increasing the number of DPUs (Data Processing Units) allocated to the AWS Glue job directly addresses the memory constraint caused by a tenfold increase in data volume. Glue ETL jobs run on Apache Spark, which distributes data processing across executors; more DPUs provide more memory and compute capacity, allowing the job to handle larger datasets without code changes.

Exam trap

The trap here is that candidates may assume memory errors always require code optimization (e.g., batching or partitioning), but the question explicitly asks for minimal code changes, making resource scaling the correct answer.

How to eliminate wrong answers

Option A is wrong because switching from Spark to Python Shell job type would reduce parallelism and memory capacity, as Python Shell runs on a single node with limited resources, making it unsuitable for large-scale data transformations. Option B is wrong because implementing batch processing with smaller file sizes would require significant code changes to split and manage files, contradicting the 'minimal code changes' requirement, and does not address the root cause of insufficient memory allocation. Option D is wrong because using Redshift Spectrum to query data directly from S3 bypasses the Glue ETL job entirely, which is a different architectural approach that does not resolve the memory error in the existing Glue job and may introduce new costs and complexity.

1301
MCQeasy

A data engineer is tasked with setting up a data pipeline that moves data from an on-premises Oracle database to Amazon S3 every hour. The network bandwidth is limited, and the engineer needs to ensure data consistency. Which AWS service should the engineer use?

A.AWS DataSync.
B.Amazon Kinesis Data Firehose.
C.S3 Transfer Acceleration.
D.AWS Database Migration Service (DMS) with change data capture (CDC).
AnswerD

DMS with change data capture replicates only incremental changes after the initial full load, sharply reducing bandwidth consumption on the constrained link. CDC also maintains transactional consistency between the Oracle source and Amazon S3, satisfying the hourly consistency requirement.

Why this answer

AWS DMS with CDC is the correct choice because it can continuously replicate ongoing changes from an on-premises Oracle database to Amazon S3 while ensuring data consistency. CDC captures only the incremental changes (inserts, updates, deletes) after an initial full load, minimizing the data transferred over limited bandwidth and maintaining transactional integrity.

Exam trap

The trap here is that candidates often confuse AWS DataSync (a file-transfer service) with database replication, or assume S3 Transfer Acceleration can solve bandwidth issues without addressing the need for change data capture and consistency from a live database.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is designed for large-scale file and object transfers between on-premises storage and AWS, not for streaming database changes from a relational database like Oracle. Option B is wrong because Amazon Kinesis Data Firehose is a streaming ingestion service for real-time data into S3, but it cannot directly connect to an on-premises Oracle database or perform change data capture. Option C is wrong because S3 Transfer Acceleration only speeds up uploads to S3 over the public internet by using AWS edge locations; it does not handle database replication, CDC, or data consistency from an on-premises source.

1302
Multi-Selectmedium

A company needs to enforce encryption at rest for all data stored in Amazon S3. Which of the following are valid methods to achieve this? (Choose TWO.)

Select 2 answers
A.Use Amazon S3 Transfer Acceleration.
B.Enable default bucket encryption using SSE-S3.
C.Enable S3 Versioning.
D.Use client-side encryption before uploading objects.
E.Use SSL/TLS for all S3 API calls.
AnswersB, D

Default bucket encryption with SSE-S3 applies AES-256 encryption automatically to every object written to the bucket, satisfying encryption at rest without per-object configuration. It is a bucket-level setting, so new uploads inherit it and existing unencrypted objects are not retroactively encrypted.

Why this answer

Option B is correct because enabling default bucket encryption with SSE-S3 causes Amazon S3 to automatically encrypt every object at rest using AES-256 with S3-managed keys, satisfying the requirement without any client changes. Option D is correct because client-side encryption encrypts the data before it ever reaches S3, so the objects are stored in encrypted form and remain protected at rest regardless of server-side settings. Option A is incorrect because S3 Transfer Acceleration only speeds up uploads over AWS edge locations and does not encrypt data at rest.

Option C is incorrect because S3 Versioning only preserves multiple object versions and does not provide encryption. Option E is incorrect because SSL/TLS protects data in transit between the client and S3, not data at rest in the bucket.

1303
MCQmedium

Refer to the exhibit. A data engineer applies the following S3 bucket policy to an S3 bucket. What does this policy enforce?

A.Denies all uploads unless SSE-S3 is used
B.Allows only SSE-S3 encrypted uploads
C.Allows any type of server-side encryption
D.Requires that all objects uploaded to the bucket be encrypted with SSE-KMS
AnswerD

The policy's Deny on s3:PutObject triggers when the s3:x-amz-server-side-encryption header is absent or does not equal aws:kms, so uploads must specify SSE-KMS encryption. This enforces the SSE-KMS requirement rather than merely AES256 or transport encryption.

Why this answer

The bucket policy uses a Deny effect with a condition that checks if the s3:x-amz-server-side-encryption header is not 'aws:kms'. This means any PutObject request that does not use SSE-KMS will be denied. Therefore, the policy enforces that all objects uploaded must be encrypted with SSE-KMS.

Option A is incorrect because the policy does not mention SSE-S3; it denies if not SSE-KMS. Option B is incorrect because it requires SSE-KMS, not SSE-S3. Option C is incorrect because the policy only allows SSE-KMS, not any type of server-side encryption.

Option D is correct.

1304
MCQeasy

A data engineer needs to store encryption keys used for protecting data in Amazon S3 and automatically rotate them every year. Which service should be used?

A.AWS KMS
B.AWS CloudHSM
C.AWS Certificate Manager
D.AWS Secrets Manager
AnswerA

AWS KMS stores customer master keys and supports automatic annual rotation, satisfying the yearly rotation constraint. S3 server-side encryption with SSE-KMS references these keys, so key material never leaves KMS and rotation is transparent to applications reading the protected objects.

Why this answer

AWS KMS is the managed service for creating, storing, and rotating encryption keys, and it supports automatic annual rotation for customer managed keys. It integrates natively with S3 SSE-KMS, making it the correct choice for key storage and yearly rotation.

Exam trap

The trap is confusing key management (KMS) with secret management (Secrets Manager) or certificate management (ACM) — only KMS handles encryption key storage and rotation.

How to eliminate wrong answers

Option B is wrong because CloudHSM provides dedicated hardware security modules for custom key management but does not offer automatic annual rotation as a built-in feature; it requires manual key management. Option C is wrong because Certificate Manager manages TLS/SSL certificates, not data encryption keys. Option D is wrong because Secrets Manager stores secrets like passwords and API keys, not encryption keys for S3 data.

1305
MCQhard

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Redshift. The migration uses a full load plus change data capture (CDC). During the CDC phase, the engineer notices that some updates are not being applied to the target Redshift tables. The DMS task logs show no errors. What is the MOST likely cause?

A.The Redshift target table does not have a primary key defined, causing DMS to skip updates.
B.The DMS replication instance is undersized, causing it to drop change events under load.
C.The source Oracle database is not configured for supplemental logging, so DMS cannot capture all column changes.
D.The DMS task is configured with 'Limited LOB mode' and large LOB columns are being truncated, causing updates to be ignored.
AnswerC

For Oracle CDC, DMS requires supplemental logging to be enabled at the database or table level. Without it, redo logs may not contain enough information to reconstruct updates, especially for columns not in the primary key. This can cause updates to be missed silently, as DMS cannot capture the full before/after image.

Why this answer

Oracle CDC in AWS DMS relies on supplemental logging to capture column-level changes. If supplemental logging is not enabled, the redo logs lack the necessary information for DMS to reconstruct updates, leading to missed changes without errors. Enabling supplemental logging at the database or table level resolves the issue.

Exam trap

The trap here is assuming that target-side constraints like primary keys or LOB settings cause missing updates, when the root cause is source-side CDC configuration.

1306
Multi-Selectmedium

A data engineer is designing a data lake on Amazon S3 that will store sensitive financial data. The engineer needs to implement encryption at rest and ensure that only authorized users can access the data. Which TWO actions should the engineer take to meet these requirements? (Choose TWO.)

Select 2 answers
A.Configure a bucket policy that denies writes if the object is not encrypted.
B.Use server-side encryption with customer-provided keys (SSE-C).
C.Enable S3 Transfer Acceleration for the bucket.
D.Enable object-level access control lists (ACLs).
E.Create IAM policies that grant least privilege access to users.
AnswersA, E

Bucket policies can enforce encryption and control access.

Why this answer

A bucket policy with a condition that denies writes if the object is not encrypted (e.g., using `s3:x-amz-server-side-encryption` or `s3:PutObject` with `aws:SecureTransport`) enforces encryption at rest at the time of upload. This ensures that all objects written to the bucket are encrypted, meeting the encryption requirement without relying on client-side behavior.

Exam trap

The trap here is that candidates often confuse encryption enforcement with encryption method selection, picking SSE-C (option B) because it sounds more secure, but the question asks for actions that ensure encryption at rest and authorized access, not a specific key management model.

1307
Multi-Selecthard

A data engineer is building a data ingestion pipeline using AWS Glue. The source is an Amazon DynamoDB table, and the target is an Amazon S3 data lake in Parquet format. The pipeline must handle large volumes and ensure exactly-once processing. Which THREE features should the engineer use together to achieve this? (Choose THREE.)

Select 3 answers
A.Use Amazon Kinesis Data Streams to capture DynamoDB Streams changes.
B.Configure the Glue job to convert data to Parquet format.
C.Use Amazon S3 Object Lambda to transform data on the fly.
D.Enable job bookmarks in the Glue job to track processed items.
E.Use DynamoDB's export to S3 feature to get a full snapshot.
AnswersB, D, E

Parquet is columnar and efficient for analytics.

Why this answer

Converting data to Parquet format is a core requirement for an S3 data lake, as Parquet offers columnar storage, compression, and efficient querying via services like Amazon Athena and Amazon Redshift Spectrum. AWS Glue natively supports Parquet as an output format, enabling the engineer to specify it in the job's output schema or transformation logic.

Exam trap

The trap here is that candidates often confuse streaming services like Kinesis Data Streams with batch processing, assuming they are required for exactly-once guarantees, when in fact AWS Glue job bookmarks combined with DynamoDB Streams or export to S3 provide a simpler and more reliable solution.

1308
MCQeasy

A company wants to enforce that all data in Amazon S3 is encrypted at rest. They want to automatically reject any PUT request that does not include encryption headers. What S3 feature should they use?

A.Bucket policy with a condition for encryption headers
B.Default encryption
C.MFA Delete
D.S3 Block Public Access
AnswerA

A bucket policy with `s3:PutObject` denied unless `s3:x-amz-server-side-encryption` is present rejects unencrypted PUT requests at the authorisation layer, satisfying the requirement to block uploads lacking encryption headers. This enforces encryption at rest without relying on client compliance or post-upload remediation.

Why this answer

An S3 bucket policy with a condition such as s3:x-amz-server-side-encryption or s3:x-amz-server-side-encryption-aws-kms-key-id can explicitly deny any PutObject request that lacks the required encryption headers. This enforces encryption at rest by rejecting unencrypted uploads at the API level, which is exactly the requirement. Default encryption alone does not reject unencrypted PUTs — it only encrypts them automatically.

Exam trap

DEA-C01 often tests whether candidates confuse default encryption (which silently encrypts but does not reject) with a bucket policy condition (which actively denies unencrypted PUTs) — the requirement to reject is the key differentiator.

How to eliminate wrong answers

Option B is wrong because default bucket encryption (SSE-S3 or SSE-KMS) transparently encrypts objects even when the client sends no encryption headers, so it does not reject unencrypted PUT requests — the opposite of the requirement. Option C is wrong because MFA Delete protects against accidental or malicious deletion of object versions; it has nothing to do with encryption enforcement. Option D is wrong because S3 Block Public Access prevents public ACLs and policies; it does not inspect or require encryption headers on PUT requests.

1309
MCQmedium

A data engineer uses AWS Glue to process data from S3. The Glue job frequently fails with 'Out of Memory' errors. The job reads several large compressed files. What is the MOST effective way to resolve this issue without changing the code?

A.Increase the number of G.1X workers or use G.2X workers
B.Convert the compressed files to uncompressed format before processing
C.Repartition the data to fewer partitions
D.Increase the job timeout setting
AnswerA

Glue workers hold the memory used for shuffles and decompression, so adding G.1X workers or switching to G.2X doubles per-worker memory and vCPU. That relieves the out-of-memory condition caused by large compressed files without editing the job script.

Why this answer

Increasing the number of G.1X workers or switching to G.2X workers directly addresses the 'Out of Memory' errors by allocating more memory per Spark executor. G.1X provides 16 GB of memory per worker, while G.2X provides 32 GB, which is critical when processing large compressed files because decompression and transformation require additional heap space. This approach resolves the issue without modifying the job code, as it only changes the resource configuration.

Exam trap

The trap here is that candidates often confuse 'Out of Memory' errors with performance issues and choose to reduce parallelism (Option C) or increase timeout (Option D), not realizing that memory exhaustion requires more memory per executor, not fewer tasks or longer runtime.

How to eliminate wrong answers

Option B is wrong because converting compressed files to uncompressed format would increase the data volume read from S3, potentially worsening memory pressure and increasing I/O costs, and it requires code changes to handle the new format. Option C is wrong because repartitioning to fewer partitions reduces parallelism, causing each Spark task to process more data, which would exacerbate memory issues rather than resolve them. Option D is wrong because increasing the job timeout setting only extends the maximum runtime before the job is killed; it does not address the underlying memory exhaustion, so the job will still fail with 'Out of Memory' errors.

1310
MCQmedium

A company is ingesting streaming data from thousands of IoT devices into Amazon Kinesis Data Streams. The data is processed by a Kinesis Data Analytics application. Recently, the application started reporting high iterator age (millisBehindLatest). Which action would BEST reduce the iterator age?

A.Decrease the data retention period of the Kinesis stream.
B.Increase the data retention period of the Kinesis stream.
C.Increase the record size limit in the Kinesis stream.
D.Increase the number of shards in the Kinesis stream.
AnswerD

Iterator age grows when consumers cannot keep pace with incoming records. Adding shards increases the stream's parallel processing capacity, letting the Kinesis Data Analytics application consume records faster and reduce millisBehindLatest, unlike scaling the application alone.

Why this answer

Increasing the number of shards increases the stream's throughput capacity, allowing the Kinesis Data Analytics application to consume data faster and reduce the iterator age (millisBehindLatest). Option A is incorrect: decreasing the data retention period does not improve processing speed; it only reduces the time data is stored. Option B is incorrect: increasing retention also does not affect processing speed.

Option C is incorrect: the record size limit is fixed (1 MB) and cannot be increased; increasing shards is the appropriate scaling action.

1311
MCQeasy

A company stores sensitive customer data in an S3 bucket. The data engineer needs to ensure that all data is encrypted at rest. Which S3 feature should be enabled?

A.S3 Versioning
B.S3 Block Public Access
C.Bucket policy requiring aws:SecureTransport
D.Default encryption
AnswerD

Default encryption applies SSE-S3 or SSE-KMS automatically to every object written to the bucket, satisfying the encryption-at-rest requirement without per-request headers. Unlike bucket policies, which only govern access, this setting enforces cryptographic protection at the storage layer itself, covering all uploads including those from unconfigured clients.

Why this answer

Default encryption ensures that all new objects written to the bucket are encrypted at rest using SSE-S3, SSE-KMS, or SSE-C. Option A is incorrect because S3 Versioning tracks object versions but does not encrypt data. Option B is incorrect because S3 Block Public Access controls public access, not encryption.

Option C is incorrect because the aws:SecureTransport condition enforces encryption in transit, not at rest.

1312
MCQeasy

A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data must be transformed in real-time and then stored in Amazon S3. Which AWS service should be used to perform the transformation?

A.AWS Glue
B.Amazon EMR
C.Amazon Kinesis Data Analytics
D.Amazon Athena
AnswerC

Amazon Kinesis Data Analytics runs Apache Flink SQL or applications directly against Kinesis Data Streams, performing stateful real-time transformation before output is delivered onward to Amazon S3. It satisfies the stem's requirement for in-flight transformation of IoT telemetry without managing servers.

Why this answer

Amazon Kinesis Data Analytics (now part of Amazon Managed Service for Apache Flink) is the correct choice because it can consume streaming data from Kinesis Data Streams, apply real-time transformations using SQL or Apache Flink, and then output the transformed data to destinations like Amazon S3. This directly meets the requirement for real-time transformation of streaming IoT data before storage.

Exam trap

The trap here is that candidates often confuse AWS Glue's batch ETL capabilities with real-time streaming, leading them to select Glue despite its lack of native support for live Kinesis Data Streams processing.

How to eliminate wrong answers

Option A is wrong because AWS Glue is a serverless ETL service designed for batch processing and cataloging data in data lakes, not for real-time stream transformations; it cannot directly process live Kinesis Data Streams with sub-second latency. Option B is wrong because Amazon EMR is a big data platform for running frameworks like Apache Spark or Hadoop on clusters, which is overkill and not optimized for lightweight, continuous real-time transformations on a single Kinesis stream; it introduces cluster management overhead and higher latency. Option D is wrong because Amazon Athena is an interactive query service for analyzing data already stored in S3 using SQL, not for transforming streaming data in motion; it cannot ingest or process live Kinesis Data Streams.

1313
MCQmedium

A company uses AWS DMS to migrate an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. After initial load, ongoing replication is set up. The replication task shows 'Task status: failed with error: The specified LSN is not available in the source database logs.' What is the most likely cause?

A.DMS does not support PostgreSQL as a source for ongoing replication.
B.The source database's network security group blocks outbound traffic to DMS.
C.The full load was incomplete, preventing CDC from starting.
D.The source database's WAL retention period is too short, and required logs have been purged.
AnswerD

PostgreSQL logical replication requires the source WAL segments containing the task's restart LSN. If wal_keep_size or the replication slot's retention is too small, those segments are recycled, so DMS cannot resume and reports the LSN as unavailable.

Why this answer

The error 'The specified LSN is not available in the source database logs' indicates that AWS DMS is trying to read a Write-Ahead Log (WAL) position that has already been purged. PostgreSQL sources require sufficient WAL retention to allow DMS to catch up during ongoing replication (CDC). If the WAL segments are recycled or removed before DMS reads them, the replication task fails with this specific LSN error.

Exam trap

The trap here is that candidates confuse connectivity issues (Option B) or general CDC support (Option A) with the specific LSN error, which is a WAL retention problem unique to PostgreSQL logical replication.

How to eliminate wrong answers

Option A is wrong because AWS DMS fully supports PostgreSQL as a source for ongoing replication using logical replication slots and WAL-based CDC. Option B is wrong because network security group blocks would cause connection timeout or 'Unable to connect' errors, not an LSN availability error. Option C is wrong because an incomplete full load would prevent the task from starting CDC at all, but the error message specifically refers to missing LSN in logs, which occurs after CDC has begun and the source WAL has been purged.

1314
MCQhard

A data pipeline uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The delivery stream is configured with a buffer size of 5 MB and a buffer interval of 60 seconds. The team notices that the S3 objects are much smaller than 5 MB. What is the most likely explanation?

A.The incoming data volume is low, so the 60-second buffer interval triggers delivery before the 5 MB buffer is filled.
B.The S3 bucket has event notifications that split the objects.
C.The S3 bucket has a lifecycle policy that transitions objects to Glacier.
D.The delivery stream is using GZIP compression, which reduces the object size.
AnswerA

Firehose delivers when either threshold is reached. With low incoming volume, the 5 MB buffer never fills, so the 60-second interval triggers delivery first, producing objects smaller than 5 MB. This directly explains the stem's observation.

Why this answer

If the incoming data rate is low, the buffer interval (60 seconds) expires before the buffer size (5 MB) is reached, causing small objects. Option B is incorrect because S3 event notifications do not split objects; they are triggered by events but do not affect object size. Option C is incorrect because S3 lifecycle policies transition objects to Glacier, which does not affect object size during delivery.

Option D is incorrect because GZIP compression reduces the object size after batching, but the buffer interval can still trigger delivery before the buffer is full, so it is not the most likely explanation.

1315
MCQeasy

Refer to the exhibit. A data engineer runs the above CLI command to find files smaller than 1000 bytes in a bucket. The command returns an empty array, but the engineer knows there are small files. What is the issue?

A.The prefix is incorrect; it should be 'logs/2023/01/01/'.
B.The bucket policy does not allow listing objects.
C.The query syntax is invalid; use a filter instead.
D.The Size is compared as a string, not an integer; remove quotes around '1000'.
AnswerD

JMESPath comparison requires numeric types.

Why this answer

In the AWS CLI `list-objects-v2` command with `--query`, the `Size` field is a numeric value, but the query string `Size < '1000'` compares it as a string. This causes a lexicographic comparison, so files with sizes like '900' would be correctly matched, but any size with more digits (e.g., '1000' itself or '999') may fail due to string ordering. Removing the quotes around `1000` treats it as an integer, enabling proper numeric comparison.

Exam trap

The DEA-C01 exam often tests the subtle distinction between string and numeric comparisons in JMESPath queries, where candidates assume quoted values are automatically coerced to numbers, but in reality, quotes force string comparison.

How to eliminate wrong answers

Option A is wrong because the prefix 'logs/2023/01/01/' is not necessarily incorrect; the engineer knows small files exist, and the prefix is just a filter—if files are under that prefix, the issue is not the prefix. Option B is wrong because if the bucket policy did not allow listing objects, the CLI command would return an access denied error, not an empty array. Option C is wrong because the query syntax using `--query` with JMESPath expressions is valid; a filter is not required for this comparison, and the syntax `Size < '1000'` is syntactically correct but semantically wrong due to type coercion.

1316
MCQmedium

A data engineer is troubleshooting a slow Amazon Redshift query that joins several large tables. The query plan shows a large number of broadcasts. Which design change would most likely reduce the broadcast operations?

A.Change the SORT KEY on all tables to match the join column.
B.Change the DISTSTYLE to EVEN on all tables.
C.Change the DISTKEY on all tables to match the join column.
D.Change the DISTSTYLE to ALL on all large tables.
AnswerC

Broadcasts occur when joined tables lack a common distribution key, forcing each node to copy rows. Setting DISTKEY to the join column co-locates matching rows on the same slice, enabling collocated joins and eliminating broadcast traffic across the cluster.

Why this answer

Setting the DISTKEY on all tables to the join column ensures that rows with the same join key value are co-located on the same compute node. This allows Redshift to perform a collocated join, eliminating the need to broadcast entire tables across the network, which is the primary cause of the slow query.

Exam trap

The trap here is that candidates confuse SORT KEY (which optimizes data skipping and range scans) with DISTKEY (which controls data distribution for joins), leading them to pick Option A, even though broadcast reduction is purely a distribution concern.

How to eliminate wrong answers

Option A is wrong because changing the SORT KEY affects the order of data on disk and can improve range-restricted scans, but it does not influence data distribution across nodes; broadcast operations are caused by distribution mismatches, not sort order. Option B is wrong because changing DISTSTYLE to EVEN distributes rows randomly across nodes, which maximizes the chance that join keys are scattered, forcing Redshift to broadcast rows to satisfy the join. Option D is wrong because changing DISTSTYLE to ALL on large tables copies the entire table to every node, which reduces broadcasts but at the cost of massive storage and maintenance overhead, making it impractical for large tables and often degrading overall performance.

1317
MCQhard

A data engineer is using AWS Glue to process a large dataset stored in Amazon S3. The dataset is partitioned by year, month, and day. The engineer notices that the Glue job is taking a long time and consuming many DPUs. The job reads all partitions, filters the data, and writes the result to another S3 location. The engineer wants to optimize the job to process only the required partitions and reduce cost. Which action should the engineer take?

A.Convert the dataset to Parquet format before processing.
B.Increase the number of DPUs to speed up the job.
C.Enable AWS Glue job bookmarks to track processed partitions.
D.Use predicate pushdown in the Glue job by specifying a push_down_predicate when reading from the AWS Glue Data Catalog.
AnswerD

Predicate pushdown allows the Glue job to filter partitions at the source based on the push_down_predicate parameter. When reading from the Data Catalog, you can specify a condition like 'year=2023 and month=01' to read only those partitions. This reduces the amount of data read and processed, improving performance and reducing DPU consumption. It is the recommended way to optimize partition pruning in Glue.

Why this answer

Predicate pushdown in AWS Glue allows the job to read only the partitions that match a specified condition, reducing data scanned and DPU usage. By using push_down_predicate when reading from the Data Catalog, the job filters at the source. Other options either do not reduce data read (bookmarks, more DPUs) or are not sufficient alone (Parquet conversion).

Exam trap

The trap here is thinking that job bookmarks or more DPUs will reduce the data read; bookmarks only avoid reprocessing, and more DPUs increase cost without reducing data volume.

1318
MCQeasy

A data engineer needs to share a dataset from an S3 bucket in Account A with users in Account B. The dataset must remain encrypted at rest with an S3-managed key. What is the MOST secure way to grant cross-account access?

A.Make the bucket public and use bucket policies to allow only Account B users.
B.Create a bucket policy that grants cross-account access to an IAM role in Account B.
C.Use S3 object ACLs to grant access to Account B's root user.
D.Use an S3 VPC endpoint to allow Account B users through private IPs.
AnswerB

A bucket policy granting access to an IAM role in Account B satisfies the cross-account requirement without exposing long-lived credentials. Account B users assume that role via AWS STS, receiving temporary credentials scoped by the trust policy. SSE-S3 encryption remains intact, since S3-managed keys need no cross-account KMS permissions, unlike SSE-KMS.

Why this answer

A bucket policy granting access to the IAM role in Account B is the recommended secure method for cross-account access to S3 objects encrypted with S3-managed keys. Option A is insecure because it grants public access. Option C is incorrect because ACLs are legacy and less secure for cross-account scenarios.

Option D is incorrect because while S3 VPC endpoints are a valid AWS feature that provides private connectivity to S3, they do not grant cross-account access; bucket policies are still required to authorize access.

1319
MCQhard

A company uses Amazon Kinesis Data Analytics for real-time anomaly detection on clickstream data. The application uses a sliding window of 1 minute. The data engineer notices that the application is producing incorrect results because late-arriving records are not being handled properly. What should the data engineer do to ensure late records are included in the window calculations?

A.Use a Kinesis Data Firehose to buffer the data and then send to Kinesis Data Analytics.
B.Increase the watermark delay in the Kinesis Data Analytics application to allow more time for late records.
C.Increase the window size from 1 minute to 2 minutes.
D.Increase the retention period of the Kinesis stream to 7 days.
AnswerB

Increasing the watermark delay lets the application tolerate out-of-order events by extending how long it waits before closing a window, so late-arriving clickstream records still fall inside the 1-minute sliding window and are included in the anomaly calculations. This directly satisfies the requirement to handle late data correctly.

Why this answer

Kinesis Data Analytics uses watermarks to track event time progress and determine when to finalize window calculations. Increasing the watermark delay allows the application to wait longer for late-arriving records before closing the window, ensuring they are included in the aggregation.

Exam trap

The trap here is that candidates confuse stream retention (how long data is stored) with watermark delay (how long the application waits for late events), leading them to incorrectly choose option D.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Firehose is a delivery stream that buffers data for loading into destinations like S3 or Redshift; it does not provide late-record handling logic for Kinesis Data Analytics windows. Option C is wrong because increasing the window size from 1 minute to 2 minutes does not address late arrivals—it simply aggregates over a longer period, which can still miss records that arrive after the window's end time. Option D is wrong because increasing the retention period of the Kinesis stream to 7 days only affects how long data is stored in the stream, not how the analytics application handles late-arriving records within its window computations.

1320
MCQhard

A data engineer is using AWS Glue to read data from an Amazon Kinesis Data Stream. The Glue job is configured to process the stream in micro-batches. The engineer notices that the job is not processing all records and sometimes skips data. The Kinesis stream has multiple shards, and the Glue job is using the 'kinesis' connection type. What is the most likely cause of the missing records?

A.The Glue job is reading from the stream using the 'GetRecords' API without specifying a shard iterator, causing it to miss records.
B.The Glue job is not checkpointing its progress correctly, so it reprocesses or skips records when the job restarts or fails.
C.The Glue job is not configured with the correct starting position, so it starts reading from the latest record instead of the earliest.
D.The Glue job is using the 'kinesis' connection but has not enabled the '--enable-auto-scaling' option, causing it to under-provision shards.
AnswerB

AWS Glue uses checkpoints to track the last processed record in a Kinesis stream. If checkpointing is not configured or fails, the job may start from an incorrect position after a restart, leading to skipped or reprocessed records. This is a common cause of missing data in stream processing jobs. Ensuring proper checkpointing is essential for exactly-once or at-least-once processing.

Why this answer

Checkpointing is critical for AWS Glue jobs reading from Kinesis Data Streams. Without proper checkpointing, the job cannot track which records have been processed, leading to data loss or duplication upon restart. The most likely cause of missing records is incorrect or disabled checkpointing, which causes the job to resume from an outdated position.

Exam trap

The trap here is focusing on starting position or scaling issues, when the core problem is often checkpointing, which ensures continuity across job runs.

1321
MCQhard

A data engineer is orchestrating a multi-step ingestion workflow where CSV files land in Amazon S3, an AWS Glue job transforms them, and the output is loaded into Amazon Redshift. The engineer wants conditional branching, retry logic, and the ability to pass parameters between steps. Which AWS service should be used to orchestrate this workflow?

A.AWS Glue workflows with triggers that chain crawlers and jobs together.
B.Amazon EventBridge rules that react to S3 events and start the Glue job.
C.AWS Step Functions with a state machine that invokes Glue and Redshift tasks.
D.AWS Lambda functions chained by direct invocations to run each pipeline step.
AnswerC

Step Functions supports conditional branching through choice states, configurable retry and catch logic, and passing data between states. It can invoke AWS Glue jobs and Amazon Redshift operations, making it suitable for orchestrating a multi-step ingestion workflow with the exact control-flow features described. This directly satisfies branching, retries, and parameter passing across the pipeline steps.

Why this answer

Step Functions is a state-machine-based orchestrator that natively supports choice states for conditional branching, retry and catch blocks for fault tolerance, and input and output passing between states. It integrates with AWS Glue and Amazon Redshift, so the entire CSV-to-transformed-to-Redshift workflow can be modeled, monitored, and retried as a single managed execution with clear visibility into each step.

Exam trap

The trap here is assuming that Glue triggers or EventBridge rules provide full workflow orchestration, when they lack rich branching and retry semantics.

Page 17

Page 18 of 18