Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 901–975

1321 questions total · 18pages · All types, answers revealed

Page 12

Page 13 of 18

Page 14
901
MCQeasy

A company needs to centralize audit logs from multiple AWS accounts into a single S3 bucket. Which service should be used to aggregate these logs?

A.AWS Config
B.Amazon Kinesis Data Firehose
C.AWS CloudTrail
D.Amazon CloudWatch Logs
AnswerC

AWS CloudTrail delivers account activity events across all accounts to one centralised S3 bucket, satisfying the multi-account aggregation requirement. Its organisation trail captures management and data events from every member account, writing them to a single destination bucket, which is precisely the centralised audit-log consolidation the stem demands.

Why this answer

AWS CloudTrail can be configured to deliver logs from multiple accounts to a single S3 bucket using a trail in the management account. Option C is correct.

902
MCQhard

A company runs an Apache Spark job on Amazon EMR that writes output to an S3 bucket. The job fails with the error 'S3AccessDeniedException' when writing the final output, but earlier stages succeed. The EMR cluster uses a service role and an instance profile. The S3 bucket policy allows access from the VPC only. What is the MOST likely cause?

A.The S3 bucket uses SSE-C encryption, and the EMR cluster does not have the encryption key.
B.The EMR service role does not have permissions to write to the S3 bucket.
C.The EMR cluster is not using a VPC endpoint for S3, so requests are denied by the bucket policy's VPC condition.
D.The S3 bucket is configured with 'Bucket owner enforced' setting for ACLs, and the EMR cluster's account is not the bucket owner.
AnswerC

The bucket policy restricts access to VPC, but since the Spark job runs on EMR, its requests originate from inside the VPC only if a VPC endpoint is used; otherwise, they come from public IPs.

Why this answer

The bucket policy restricts access to requests originating from the VPC, typically using a condition like `aws:SourceVpc`. If the EMR cluster does not use a VPC endpoint for S3 (either Gateway or Interface endpoint), traffic from the cluster to S3 traverses the public internet and does not match the VPC condition, causing the `S3AccessDeniedException`. Earlier stages may succeed if they use cached data or different paths, but the final write fails because it hits the bucket policy check.

Exam trap

The trap here is that candidates often assume the EMR service role (EMR_EC2_DefaultRole) is responsible for all S3 access, but in reality the instance profile (EC2 instance role) handles data plane operations, and the bucket policy's VPC condition is the key blocker when earlier stages succeed but final writes fail.

How to eliminate wrong answers

Option A is wrong because SSE-C encryption requires the client to provide the encryption key; if the key were missing, the error would be an encryption-related error (e.g., 'InvalidArgument' or 'AccessDenied' with a different message), not a generic 'S3AccessDeniedException'. Option B is wrong because the EMR service role is used for the cluster's service-level permissions (e.g., launching instances, reading logs), not for data access to S3; the instance profile (IAM role attached to EC2 instances) handles data read/write permissions, and the question states earlier stages succeed, indicating the instance profile has write permissions. Option D is wrong because the 'Bucket owner enforced' setting (S3 Object Ownership) controls ACLs and ownership of objects, not access permissions; it does not cause an 'S3AccessDeniedException' — it would affect who owns new objects, not whether the write is allowed.

903
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs and Amazon EMR steps. The state machine has a task that starts a Glue job and waits for completion using the 'StartJobRun' API with the 'sync' integration. Occasionally, the Step Functions execution fails with the error: 'States.TaskFailed: Glue job failed with error: ResourceNumberLimitExceededException'. The engineer confirms that the Glue job itself runs successfully when triggered manually. What is the most likely cause of this intermittent failure?

A.The Glue job's script has a bug that causes it to fail only when run in parallel.
B.The IAM role used by Step Functions lacks permissions to start the Glue job.
C.The AWS Glue job is hitting the concurrent job run limit for the account.
D.The Step Functions state machine is exceeding the maximum number of concurrent executions.
AnswerC

AWS Glue enforces a limit on the number of concurrent job runs per account. When multiple Step Functions executions trigger the same Glue job simultaneously, or when other processes start Glue jobs, the account may exceed the concurrent runs quota, causing the 'ResourceNumberLimitExceededException'. Since the job succeeds manually, the issue is likely contention for Glue concurrent runs, not the job code.

Why this answer

The error 'ResourceNumberLimitExceededException' from AWS Glue indicates that the account has reached the maximum number of concurrent job runs. Step Functions may trigger the job multiple times concurrently, or other processes may be running Glue jobs. The engineer should check the Glue service quotas and either increase the concurrent runs limit or implement a queueing mechanism to serialize job starts.

Exam trap

The trap here is focusing on Step Functions limits, but the error explicitly comes from AWS Glue and relates to its concurrent job runs quota.

904
MCQmedium

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon S3 in Parquet format. The source data is in JSON format and contains nested structures. The engineer needs to flatten the nested data and write it to Parquet. Which AWS Glue transform should the engineer use to flatten the nested structure?

A.Map
B.Relationalize
C.ApplyMapping
D.Filter
AnswerB

The Relationalize transform in AWS Glue flattens nested JSON structures into a relational schema, producing multiple tables that can be joined. It is specifically designed to handle nested data and is the correct choice for flattening. It preserves the relationships between the flattened tables through generated keys.

Why this answer

The Relationalize transform in AWS Glue is designed to flatten nested JSON structures into a relational format. It automatically creates multiple DynamicFrames for nested arrays and objects, generating keys to join them. This makes it the correct choice for the scenario, as it simplifies the flattening process without custom code.

Exam trap

The trap here is assuming that a generic transform like Map can flatten nested data with custom code, but Relationalize is the specialized transform for this purpose.

905
MCQmedium

A data engineer runs an AWS Glue ETL job that joins a 4 TB Parquet dataset in Amazon S3 with a small 40 MB reference lookup table stored as a single CSV file in S3. The join is taking hours and the job frequently fails with executor out-of-memory errors. The engineer wants to reduce shuffle and memory pressure with the least development effort. Which approach should the engineer take?

A.Increase the number of Glue DPU workers so more executors can hold the shuffled join partitions.
B.Convert the small lookup CSV into a broadcast join by loading it with a broadcast hint, avoiding a shuffle of the large dataset.
C.Repartition the large Parquet dataset by the join key using a Glue transform before the join.
D.Enable the AWS Glue job bookmark and rerun the job so only new files are processed.
AnswerB

Broadcasting the 40 MB lookup table ships it to every executor so the large 4 TB dataset never has to be shuffled for the join. This eliminates the expensive shuffle and the memory pressure caused by co-locating both sides, directly addressing the OOM failures. It requires only a small change to the join logic, satisfying the least-effort requirement.

Why this answer

Because one side of the join is only 40 MB, broadcasting that lookup table to every executor removes the need to shuffle the multi-terabyte dataset, cutting network I/O and eliminating the memory spikes that cause executor OOM failures. This is a minimal code change, satisfying the least-effort constraint, whereas scaling workers or repartitioning the large side still incurs a full shuffle.

Exam trap

The trap here is assuming that adding more DPU workers will fix a shuffle-heavy join, when the real fix is to change the join strategy so the large dataset is never shuffled.

906
MCQeasy

A data engineer needs to schedule a recurring AWS Glue ETL job to run every day at 2:00 AM UTC. The job must be triggered automatically without manual intervention. Which AWS service should the engineer use to create this schedule?

A.AWS Glue triggers
B.Amazon EventBridge (CloudWatch Events) rules
C.AWS Lambda with a CloudWatch Events schedule
D.AWS Step Functions with a Wait state
AnswerA

AWS Glue triggers are designed to start jobs, crawlers, or other triggers based on a schedule or events. A time-based trigger can be configured with a cron expression to run daily at 2:00 AM UTC. This is the native and simplest way to schedule Glue jobs, requiring no additional services. It integrates directly with the Glue job and provides status monitoring.

Why this answer

AWS Glue triggers provide a built-in scheduling capability for Glue jobs. A time-based trigger uses a cron expression to define the schedule, such as 'cron(0 2 * * ? *)' for 2:00 AM UTC daily. This eliminates the need for external services and keeps the scheduling logic within AWS Glue, simplifying management and monitoring.

Exam trap

The trap here is overcomplicating the solution by involving additional services like EventBridge or Lambda, when AWS Glue triggers natively support scheduled execution.

907
MCQeasy

A data engineer needs to ensure that data in transit between an Amazon RDS for PostgreSQL database and an application is encrypted. Which configuration should be used?

A.Use VPC peering to connect the application to the database
B.Enable SSL/TLS for the database connection
C.Enable encryption at rest for the RDS instance
D.Use IAM database authentication
AnswerB

SSL/TLS encrypts the wire protocol between the application and PostgreSQL, protecting credentials and query results from interception. Enabling it on the RDS connection satisfies the requirement for encryption of data in transit rather than at rest.

Why this answer

Enabling SSL/TLS for the database connection encrypts the data in transit between the application and the RDS PostgreSQL instance, ensuring confidentiality. Option A (VPC peering) provides network connectivity but does not encrypt traffic. Option C (encryption at rest) protects data stored on disk, not in transit.

Option D (IAM database authentication) controls access but does not encrypt the connection.

908
MCQmedium

A company wants to enable automatic encryption for all new objects written to an S3 bucket. The bucket has existing objects that are unencrypted. Which solution meets these requirements with the least operational overhead?

A.Configure a lifecycle policy to transition objects to a new bucket with encryption
B.Enable default encryption on the bucket using SSE-S3
C.Use S3 server-side encryption with S3 managed keys (SSE-S3) and apply a bucket policy that denies writes without encryption
D.Use S3 Batch Operations to copy existing objects with SSE-S3
AnswerB

Enabling SSE-S3 default encryption applies AES-256 encryption automatically to every new object with no key management, and existing unencrypted objects remain readable. This satisfies both requirements with minimal operational overhead compared with SSE-KMS or re-uploading objects.

Why this answer

Enabling default encryption on the bucket with SSE-S3 causes all new objects written to the bucket to be encrypted automatically without any client-side changes or bucket policy enforcement. It is a single bucket-level setting, so it has the least operational overhead. Existing unencrypted objects remain unencrypted until rewritten, but the requirement only specifies automatic encryption for new objects.

Exam trap

The trap is reading 'existing objects are unencrypted' as a requirement to encrypt them — the question only requires automatic encryption for new objects, so candidates who over-engineer with Batch Operations or bucket policies pick the wrong answer.

How to eliminate wrong answers

Option A is wrong because a lifecycle policy transitions objects between storage classes or expires them — it cannot encrypt objects, and moving to a new bucket does not retroactively encrypt existing data. Option C is wrong because while a bucket policy denying unencrypted PUTs enforces encryption, it requires clients to send encryption headers and can break existing writers; it is more operational overhead than simply enabling default encryption. Option D is wrong because S3 Batch Operations to copy existing objects with SSE-S3 addresses existing objects, not new writes, and adds operational overhead for a requirement that only concerns new objects.

909
MCQeasy

A data engineer needs to automate the backup of an Amazon RDS for PostgreSQL database. Which AWS service can be used to schedule and manage the backups?

A.Amazon S3
B.AWS Lambda
C.AWS Backup
D.Amazon CloudWatch
AnswerC

AWS Backup provides a centralised, policy-driven service that schedules and manages backups across supported resources, including Amazon RDS for PostgreSQL. This satisfies the stem's requirement to automate scheduling and management, rather than relying on manual snapshots or RDS automated backup windows alone.

Why this answer

AWS Backup is a fully managed backup service that can automate backups of RDS databases. Option A (Amazon S3) is incorrect because S3 is an object storage service, not a backup scheduling service. Option B (AWS Lambda) is incorrect because while Lambda can be used for custom automation, it is not the primary managed service for backup scheduling.

Option D (Amazon CloudWatch) is incorrect because CloudWatch is for monitoring and observability, not backup management.

910
MCQeasy

A company uses AWS DMS to replicate data from an on-premises Oracle database to Amazon RDS for MySQL. The full load completes successfully, but ongoing replication (CDC) is failing with a 'Failed to add supplemental logging' error. What should the data engineer do to resolve this issue?

A.Enable supplemental logging on the source Oracle database manually.
B.Recreate the DMS endpoint for the source database with a new connection.
C.Modify the target MySQL database to use a different engine version.
D.Increase the task log interval in the DMS task settings.
AnswerA

Oracle requires supplemental logging to supply the primary key and before-image data DMS needs for change capture. Enabling it manually on the source tables or database lets CDC read redo logs and clear the error.

Why this answer

AWS DMS requires supplemental logging to be enabled on the source Oracle database to capture change data for CDC replication. The 'Failed to add supplemental logging' error occurs because DMS cannot automatically enable this at the database level when the user lacks the necessary privileges (e.g., ALTER DATABASE). The data engineer must manually enable supplemental logging on the source Oracle database, typically by executing ALTER DATABASE ADD SUPPLEMENTAL LOG DATA; and ensuring table-level supplemental logging is also configured.

Exam trap

DEA-C01 often tests the specific prerequisite that AWS DMS requires supplemental logging on Oracle sources for CDC, and candidates may incorrectly assume DMS can always enable it automatically or that the error relates to endpoint or network configuration.

How to eliminate wrong answers

Option B is wrong because recreating the DMS endpoint does not address the missing supplemental logging configuration on the source database; the endpoint connection is already functional since the full load succeeded. Option C is wrong because modifying the target MySQL engine version is irrelevant to the source Oracle supplemental logging requirement. Option D is wrong because increasing the task log interval only affects logging verbosity and does not resolve the underlying permission or configuration issue with supplemental logging.

911
MCQeasy

Refer to the exhibit. An S3 event notification is configured to trigger an AWS Lambda function when objects are created in 'my-bucket'. The Lambda function processes the JSON file and writes results to Amazon DynamoDB. The function fails with a timeout error. Which action should the engineer take to resolve the issue?

A.Modify the S3 event notification to use a different event type
B.Grant the Lambda function permission to access DynamoDB
C.Change the trigger to Amazon SQS instead of S3
D.Increase the Lambda function timeout
AnswerD

The Lambda function exceeds its configured execution limit while processing the JSON and writing to DynamoDB, so raising the timeout lets the invocation complete. This addresses the timeout error directly, unlike memory, concurrency or retry adjustments that leave the duration ceiling unchanged.

Why this answer

The Lambda function is failing with a timeout error, which indicates that the function is taking longer to execute than the default timeout of 3 seconds. Increasing the Lambda function timeout allows the function to run longer and complete its processing of the JSON file and DynamoDB write operation without being prematurely terminated.

Exam trap

The DEA-C01 exam often tests the distinction between timeout errors and permission errors, leading candidates to incorrectly choose a permissions fix (Option B) when the error message explicitly states 'timeout'.

How to eliminate wrong answers

Option A is wrong because changing the S3 event notification to a different event type (e.g., from s3:ObjectCreated:* to s3:ObjectCreated:Put) does not address the timeout issue; the function still fails due to execution duration, not the trigger event. Option B is wrong because a timeout error is not a permissions issue; if the Lambda function lacked DynamoDB permissions, it would fail with an access denied error (e.g., 403), not a timeout. Option C is wrong because changing the trigger to Amazon SQS instead of S3 does not resolve the timeout; the Lambda function would still have the same execution duration limit and would timeout regardless of the trigger source.

912
MCQmedium

A data engineer is designing a pipeline to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real-time with minimal latency, and the engineer wants to use a fully managed service that automatically scales. The data is in JSON format and needs to be converted to Parquet before storage. Which AWS service should be used to achieve this?

A.Amazon Kinesis Data Firehose with record format conversion enabled to Parquet.
B.AWS Lambda function that reads from the Kinesis stream, converts JSON to Parquet, and writes to S3.
C.AWS Glue streaming ETL job that reads from Kinesis, converts to Parquet, and writes to S3.
D.Amazon Kinesis Data Analytics application that reads from Kinesis and writes Parquet to S3.
AnswerA

This option is correct because Amazon Kinesis Data Firehose is a fully managed service that can ingest data from a Kinesis Data Stream and deliver it to Amazon S3. It supports record format conversion from JSON to Parquet using an AWS Glue Data Catalog table schema. Firehose automatically scales to handle the throughput and provides near real-time delivery (with buffer intervals as low as 60 seconds). This meets all requirements without custom code.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is a fully managed service that can ingest from a Kinesis Data Stream, automatically scale, and deliver to Amazon S3 with near real-time latency. It supports record format conversion from JSON to Parquet using a Glue Data Catalog table, eliminating the need for custom code.

Exam trap

The trap here is overlooking Firehose's built-in format conversion capability and assuming that a custom Lambda or Glue job is required to convert JSON to Parquet.

913
MCQhard

A data engineer is using AWS Lake Formation to manage access to a data lake stored in Amazon S3. The engineer grants a data analyst SELECT permission on a table in the AWS Glue Data Catalog. However, when the analyst queries the table using Amazon Athena, they receive an error that they are not authorized to access the underlying S3 data. What is the MOST likely reason?

A.The AWS Glue Data Catalog table is encrypted with SSE-KMS and the analyst lacks kms:Decrypt.
B.The analyst's IAM role lacks the lakeformation:GetDataAccess permission.
C.The analyst's IAM role does not have s3:GetObject permission on the S3 bucket.
D.The S3 bucket is not registered as a data location in Lake Formation.
AnswerD

Lake Formation requires that the S3 location be registered as a data location to manage fine-grained access. Without registration, Lake Formation cannot grant access to the underlying data, even if table permissions are granted. The analyst would need direct S3 permissions or the location must be registered.

Why this answer

For Lake Formation to manage access to data lake data, the S3 bucket must be registered as a data location. Without registration, Lake Formation cannot enforce or grant access to the underlying objects, resulting in authorization errors even when table permissions are granted. Registering the location enables Lake Formation to provide temporary credentials for data access.

Exam trap

The trap here is assuming that granting table permissions in Lake Formation automatically grants access to the underlying S3 data, overlooking the need to register the data location.

914
MCQmedium

A company has an Amazon RDS for MySQL DB instance with read replicas. The primary DB instance fails. What is the correct procedure to promote a read replica to become the new primary?

A.Modify the read replica to be a Multi-AZ deployment and failover will occur.
B.RDS automatically fails over to the read replica within 5 minutes.
C.Manually promote the read replica to a standalone DB instance.
D.Delete the primary and the read replica will automatically become the primary.
AnswerC

Promoting a read replica manually converts it into a standalone DB instance, which is the only supported route when the primary fails without Multi-AZ. The stem specifies RDS for MySQL with read replicas, so no automatic failover exists; you must trigger promotion yourself to restore write capability.

Why this answer

When an Amazon RDS for MySQL primary DB instance fails, read replicas do not automatically become the new primary. The correct procedure is to manually promote the read replica using the AWS Management Console, CLI, or API, which converts it into a standalone DB instance. After promotion, you must update your application endpoints to point to the new primary, as RDS does not handle this automatically.

Exam trap

The trap here is that candidates confuse read replicas with Multi-AZ standby instances, assuming automatic failover applies to both, but RDS read replicas require manual promotion and do not provide automatic failover.

How to eliminate wrong answers

Option A is wrong because modifying a read replica to be Multi-AZ does not trigger a failover; Multi-AZ is a separate feature for high availability within a single region, and read replicas are not part of the Multi-AZ failover mechanism. Option B is wrong because RDS does not automatically fail over to a read replica; automatic failover only occurs with Multi-AZ deployments, not with read replicas. Option D is wrong because deleting the primary DB instance does not cause the read replica to automatically become the primary; the read replica remains a read-only copy until manually promoted.

915
MCQmedium

A data engineering team is using Amazon EMR to process large datasets stored in Amazon S3. The cluster uses Spot Instances for cost savings. During processing, the team notices that tasks are failing due to Spot Instance interruptions. The team needs to make the EMR job resilient to Spot interruptions without increasing costs significantly. Which solution should they implement?

A.Use EMR instance fleets with a mix of Spot and On-Demand, but set the allocation strategy to 'lowest price'.
B.Increase the number of core nodes using On-Demand instances.
C.Use only Spot Instances but enable automatic termination and checkpointing.
D.Use EMR instance fleets with a mix of Spot and On-Demand, setting the allocation strategy to 'diversified' and using On-Demand for core nodes.
AnswerD

Instance fleets with a diversified allocation strategy spread Spot capacity across multiple instance types and Availability Zones, reducing simultaneous interruption risk, while On-Demand core nodes preserve HDFS and task durability. This keeps costs low without losing resilience.

Why this answer

Using EMR instance fleets with a mix of Spot and On-Demand instances, setting the allocation strategy to 'diversified' and using On-Demand for core nodes, provides resilience against Spot interruptions while controlling costs. The 'diversified' strategy spreads Spot instances across multiple instance pools and Availability Zones, reducing the impact of a single Spot interruption. Core nodes, which handle HDFS data and run tasks, are kept on On-Demand to ensure stability, while task nodes can use Spot for cost savings.

Exam trap

The trap is assuming that any mix of Spot and On-Demand is sufficient, but the allocation strategy and node type assignment are critical: using 'lowest price' or putting core nodes on Spot can lead to frequent job failures despite having On-Demand instances.

How to eliminate wrong answers

Option A is wrong because the 'lowest price' allocation strategy concentrates Spot instances in the cheapest pool, which increases the risk of widespread interruptions if that pool's Spot capacity is reclaimed. Option B is wrong because increasing core nodes with On-Demand instances significantly raises costs and does not address the resilience of task nodes that may still be on Spot. Option C is wrong because using only Spot Instances, even with checkpointing, does not prevent job failures from interruptions; checkpointing helps resume but does not eliminate the need for On-Demand capacity for critical nodes.

Option D is correct because it balances cost and resilience by diversifying Spot instances and using On-Demand for core nodes.

916
MCQhard

A data engineer is using AWS Lake Formation to manage access to a data lake stored in Amazon S3. The engineer needs to grant a data analyst read access to specific columns in a table registered in the AWS Glue Data Catalog, while hiding other columns that contain personally identifiable information. The analyst uses Amazon Athena to query the table. Which Lake Formation feature should the engineer use?

A.Use AWS Glue Studio to create an ETL job that copies only the allowed columns to a new table and grant access to that table.
B.Use Lake Formation column-level permissions to grant SELECT on only the allowed columns.
C.Create a resource link to the table and grant SELECT on the link.
D.Apply an IAM policy to the analyst that allows s3:GetObject on the S3 path of the table.
AnswerB

Lake Formation supports fine-grained access control, including column-level permissions. You can grant SELECT on a subset of columns in a table. When the analyst queries via Athena, Lake Formation filters the columns, and the analyst can only see the permitted columns. This directly meets the requirement to hide PII columns while allowing access to others.

Why this answer

Lake Formation column-level permissions allow granular control over which columns a principal can access. By granting SELECT only on specific columns, the analyst can query the table via Athena but will receive errors or filtered results for unauthorized columns. This is the intended feature for hiding PII while enabling access to other data.

Exam trap

The trap here is confusing resource links with column-level permissions; resource links alone do not provide column filtering.

917
MCQmedium

Refer to the exhibit. The S3 bucket policy above is applied to the bucket "example-bucket". An IAM user attempts to upload an object to the bucket without specifying any encryption header. What is the outcome?

A.The upload succeeds but the object is not encrypted
B.The upload fails because GetObject requires encryption
C.The object is uploaded successfully with SSE-S3 encryption by default
D.The upload fails with an Access Denied error
AnswerD

The bucket policy denies `s3:PutObject` unless the request includes the `s3:x-amz-server-side-encryption` header, so an upload without any encryption header fails the condition and is explicitly denied. Because an explicit Deny in a bucket policy overrides any IAM allow, S3 returns Access Denied.

Why this answer

The bucket policy shown requires that any PutObject request include server-side encryption headers (e.g., x-amz-server-side-encryption). Since the IAM user uploads without specifying an encryption header, the condition in the policy is not met, and the request is denied with an Access Denied error. This is the standard behavior of a deny-by-default encryption-enforcement policy.

Exam trap

The trap is assuming S3 default encryption will silently encrypt the object; DEA-C01 tests that an explicit encryption-required bucket policy causes Access Denied when the header is missing.

How to eliminate wrong answers

Option A is wrong because the policy explicitly blocks unencrypted uploads — the object would not be stored unencrypted. Option B is wrong because GetObject is not the operation being attempted; the failure is on PutObject, and GetObject encryption requirements are a separate policy concern. Option C is wrong because S3 does not automatically apply SSE-S3 when a policy requires an explicit encryption header — the request fails before default encryption is considered.

918
MCQmedium

A company is migrating an on-premises Apache Cassandra database to Amazon Keyspaces. The database has a table with a partition key of 'user_id' and a clustering column of 'timestamp'. The application frequently queries the last 10 records for a given user. Which table design in Keyspaces would provide the BEST query performance for this access pattern?

A.Partition key: random column, clustering column: none.
B.Partition key: timestamp, clustering column: user_id.
C.Partition key: user_id, clustering column: none.
D.Partition key: user_id, clustering column: timestamp (descending order).
AnswerD

Descending clustering order stores rows on disk newest-first, so retrieving the last 10 records for a user reads one contiguous slice from the start of the partition rather than scanning the entire partition and sorting. This directly satisfies the frequent "last 10 records per user" access pattern with a single efficient query.

Why this answer

It preserves the original Cassandra table design with 'user_id' as the partition key and 'timestamp' as the clustering column in descending order. This allows Keyspaces to efficiently retrieve the last 10 records for a given user by performing a range query on the clustering column within a single partition, avoiding full table scans or cross-partition queries.

Exam trap

The trap here is that candidates may think a random partition key (Option A) or timestamp-based partition key (Option B) improves write distribution, but they overlook that the query pattern requires efficient reads within a single partition, which is best achieved by using the query filter column as the partition key and the sort column as the clustering key with the appropriate order.

How to eliminate wrong answers

Option A is wrong because using a random partition key with no clustering column would scatter data across partitions, requiring a full scan to find records for a specific user, which is highly inefficient. Option B is wrong because using 'timestamp' as the partition key would place each timestamp in a separate partition, making it impossible to query all records for a user without scanning multiple partitions, and the clustering column 'user_id' would not help retrieve the last 10 records per user efficiently. Option C is wrong because while 'user_id' as the partition key correctly groups data by user, having no clustering column means you cannot order records by timestamp, so retrieving the last 10 records would require fetching all records for that user and sorting them in application code, which is suboptimal.

919
MCQmedium

A data engineer sees this AWS Glue table definition in the Data Catalog. The engineer wants to query this table with Amazon Athena, but the query returns zero rows. What is the MOST likely cause?

A.The data files are not in the specified S3 location.
B.The SerDe library is incorrect for CSV files.
C.The table format CSV is not supported by Athena.
D.Athena cannot read tables from the Glue Data Catalog.
AnswerA

Athena reads the table's declared S3 location from the Data Catalog; if no objects exist at that prefix, the query scans nothing and returns zero rows. The partition metadata and schema may be valid, but the underlying files must reside at the specified location.

Why this answer

The most likely cause is that the data files are not in the specified S3 location. When an AWS Glue table is defined in the Data Catalog, Athena reads the table's metadata (including the S3 location) and then attempts to read the underlying data files from that exact path. If the files are missing, misnamed, or in a different prefix, Athena returns zero rows because there is no data to scan.

This is a common misconfiguration when the S3 path in the table definition does not match the actual data storage.

Exam trap

The trap here is that candidates often assume the issue is with the SerDe or format compatibility, but the most common real-world cause is simply that the data files are not present at the specified S3 location, leading to zero rows returned.

How to eliminate wrong answers

Option B is wrong because the SerDe library is not incorrect for CSV files; Athena uses the LazySimpleSerDe by default for CSV, which is fully supported and does not cause zero rows. Option C is wrong because CSV is a widely supported table format in Athena, and Athena can query CSV files natively. Option D is wrong because Athena is designed to read tables from the Glue Data Catalog; in fact, Athena and Glue Data Catalog are tightly integrated, and this is a standard use case.

920
MCQmedium

A company is using Amazon Athena to query data stored in S3. Queries are failing with 'HIVE_INVALID_PARTITION' errors. What is the most likely cause?

A.The S3 bucket is configured with a bucket policy that denies access to the Athena service.
B.A partition folder in S3 has been deleted or moved, but the table metadata still references it.
C.The data is compressed with gzip, but the table definition expects uncompressed data.
D.The data files are in CSV format but the table definition expects Parquet.
AnswerB

Athena derives partitions from S3 prefixes registered in the AWS Glue Data Catalog. If a partition folder is deleted or moved while its metadata entry remains, queries referencing that partition fail with HIVE_INVALID_PARTITION. Running MSCK REPAIR TABLE or dropping the stale partition resolves the mismatch.

Why this answer

The 'HIVE_INVALID_PARTITION' error in Amazon Athena occurs when the table's partition metadata in the AWS Glue Data Catalog (or Hive metastore) references a partition folder that no longer exists in the S3 bucket. Athena relies on the metadata to locate data files; if a partition folder is deleted or moved without updating the metadata, queries fail because Athena cannot find the expected data location.

Exam trap

The trap here is that candidates confuse permission errors (like S3 bucket policies) with metadata consistency errors, or assume compression or format mismatches cause partition-specific errors, when in reality 'HIVE_INVALID_PARTITION' is a direct indicator of a stale or missing partition folder in the catalog.

How to eliminate wrong answers

Option A is wrong because a bucket policy denying Athena access would cause an 'Access Denied' error, not a 'HIVE_INVALID_PARTITION' error, which is specific to partition metadata mismatch. Option C is wrong because Athena supports reading gzip-compressed data transparently, and compression mismatch does not produce partition-related errors. Option D is wrong because a schema mismatch between CSV and Parquet would cause a 'HIVE_CANNOT_OPEN_SPLIT' or data type conversion error, not a partition validation error.

921
MCQeasy

A data engineer needs to ingest JSON files from an S3 bucket into a DynamoDB table. The files are updated hourly and contain new records. Which AWS service should be used to trigger a Lambda function for each new object?

A.Kinesis Data Firehose
B.Amazon EventBridge
C.S3 Event Notifications
D.Amazon SQS
AnswerC

S3 Event Notifications publish an event to Lambda whenever a new object lands in the bucket, invoking the function per object. This satisfies the stem's requirement to trigger processing for each hourly-uploaded file without polling, unlike CloudWatch Events schedules or manual invocation.

Why this answer

S3 Event Notifications are the correct choice because they are designed to trigger AWS Lambda functions directly in response to object creation events (e.g., `s3:ObjectCreated:*`) in an S3 bucket. This allows the data engineer to automatically invoke a Lambda function for each new JSON file as it is uploaded, enabling ingestion into DynamoDB without polling or additional infrastructure.

Exam trap

The trap here is that candidates may confuse Amazon EventBridge with S3 Event Notifications, but EventBridge is not the native S3 event trigger—S3 Event Notifications are the direct, service-integrated mechanism for invoking Lambda on object creation.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Firehose is a streaming ingestion service that loads data into destinations like S3, Redshift, or Elasticsearch, but it cannot directly trigger a Lambda function for each new object; it is not an event source for Lambda. Option B is wrong because Amazon EventBridge can schedule or react to events from various AWS services, but it is not the native, direct integration for S3 object creation events—S3 Event Notifications are the simpler and more appropriate mechanism. Option D is wrong because Amazon SQS is a message queue service that decouples components, but it does not natively trigger Lambda from S3 events without additional configuration (e.g., S3 Event Notification to SQS, then Lambda polling SQS), which adds unnecessary complexity for this use case.

922
MCQmedium

A company uses Amazon S3 to store sensitive data. The security team wants to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted at rest using server-side encryption with AWS KMS managed keys (SSE-KMS). Which bucket policy statement should be added to enforce this requirement?

A.Deny put requests where 's3:x-amz-server-side-encryption' is 'aws:kms'
B.Deny put requests where 's3:x-amz-server-side-encryption' is not 'aws:kms'
C.Deny put requests where 's3:x-amz-server-side-encryption' is not 'AES256'
D.Deny put requests where 's3:x-amz-server-side-encryption' is not set
AnswerB

A bucket policy denying `s3:PutObject` when `s3:x-amz-server-side-encryption` is absent or not `aws:kms` enforces SSE-KMS at upload time, satisfying the requirement that every object be encrypted with AWS KMS managed keys. The condition key inspects the request header, blocking non-compliant uploads before they succeed.

Why this answer

It denies any S3 PUT request that does not include the `x-amz-server-side-encryption` header set to `aws:kms`, thereby enforcing SSE-KMS encryption for all objects uploaded to the bucket. This bucket policy condition ensures that only requests specifying AWS KMS-managed keys are allowed, meeting the security team's requirement for automatic encryption at rest with SSE-KMS.

Exam trap

The DEA-C01 exam often tests the distinction between enforcing a specific encryption type (SSE-KMS) versus simply requiring encryption (any type), so candidates may incorrectly choose Option D (deny if not set) or Option C (deny if not AES256) because they confuse 'encryption at rest' with 'SSE-KMS specifically'.

How to eliminate wrong answers

Option A is wrong because it denies PUT requests where `s3:x-amz-server-side-encryption` is `aws:kms`, which would block the very encryption method required, making it impossible to upload objects with SSE-KMS. Option C is wrong because it denies PUT requests where encryption is not `AES256`, which would enforce SSE-S3 (AES256) instead of SSE-KMS, failing the requirement for KMS-managed keys. Option D is wrong because it denies PUT requests where the encryption header is not set, which would block unencrypted uploads but does not specifically enforce SSE-KMS; it would also allow SSE-S3 or other encryption types if the header is present, missing the specific KMS requirement.

923
MCQeasy

A data engineer needs to ensure that data stored in Amazon S3 is automatically deleted after 30 days. Which S3 feature should be used?

A.S3 Lifecycle policy
B.S3 MFA Delete
C.S3 Versioning
D.S3 Object Lock
AnswerA

An S3 Lifecycle policy defines expiration rules that automatically delete objects after a specified number of days, directly meeting the 30-day deletion requirement. It applies at bucket or prefix level without custom code, unlike versioning or replication, which retain data rather than remove it.

Why this answer

S3 Lifecycle policies can automatically delete objects after a specified time period, such as 30 days. Option B (MFA Delete) requires multi-factor authentication for deletion but does not automate deletion. Option C (Versioning) keeps multiple versions but does not delete.

Option D (Object Lock) prevents deletion or modification but does not schedule automatic deletion.

924
MCQeasy

A data engineer is troubleshooting an AWS Glue ETL job that fails with the error: 'An error occurred while calling o137.pyWriteDynamicFrame. No such file or directory: s3://bucket/output/part-00000.parquet'. The job reads from a JDBC source and writes to S3. What is the most likely cause?

A.The schema of the source data has changed, causing a mismatch during write
B.The output S3 path does not exist and the Glue job does not have permission to create it
C.The Glue job ran out of memory during the transformation phase
D.The IAM role attached to the Glue job lacks permissions to read from the JDBC source
AnswerB

The error message indicates missing directory; Glue may not auto-create if permissions are insufficient.

Why this answer

The error 'No such file or directory: s3://bucket/output/part-00000.parquet' indicates that the Glue job is trying to write to an S3 path that does not exist. By default, AWS Glue does not automatically create the target S3 bucket or prefix; it requires the path to already exist or the IAM role to have s3:PutObject permissions that allow the S3 service to create the object. Since the error occurs at the write stage (pyWriteDynamicFrame) and not during read, the most likely cause is that the output S3 path does not exist and the Glue job lacks the necessary permissions to create it.

Exam trap

The trap here is that candidates often confuse a missing S3 path with a permissions issue, but the error message explicitly states 'No such file or directory', which points to the path not existing rather than a generic access denied, and the exam expects you to recognize that Glue does not auto-create the output S3 prefix.

How to eliminate wrong answers

Option A is wrong because a schema mismatch during write would typically cause a different error, such as a schema compatibility or type conversion error, not a 'No such file or directory' file system error. Option C is wrong because an out-of-memory error would manifest as a Java heap space or memory limit exception, not a missing file error on S3. Option D is wrong because the error occurs during the write phase (pyWriteDynamicFrame) and not during the JDBC read; if the IAM role lacked JDBC read permissions, the job would fail earlier with a connection or authentication error.

925
MCQeasy

A company is using Amazon DynamoDB for a gaming application. They want to store player session data that expires after 24 hours. Which DynamoDB feature should be used?

A.Time to Live (TTL)
B.DynamoDB Streams
C.Global Tables
D.Point-in-Time Recovery
AnswerA

Time to Live (TTL) lets you define an epoch-timestamp attribute so DynamoDB automatically deletes expired items at no write cost, satisfying the 24-hour session expiry without custom cleanup jobs or scans. Deletion typically occurs within 48 hours of expiry, though expired items are already hidden from reads.

Why this answer

Amazon DynamoDB Time to Live (TTL) allows you to define a per-item timestamp attribute that automatically deletes items after a specified duration. For the gaming session data that must expire after 24 hours, you can set the TTL attribute to the current time plus 24 hours, and DynamoDB will asynchronously delete expired items without any additional cost or write operations.

Exam trap

The trap here is that candidates may confuse DynamoDB Streams (which can react to deletions) with the actual mechanism that performs the deletion, or assume Point-in-Time Recovery can be used to 'roll back' expired data, neither of which addresses automatic expiration.

How to eliminate wrong answers

Option B (DynamoDB Streams) is wrong because it captures a time-ordered sequence of item-level changes (inserts, updates, deletes) in a DynamoDB table, but it does not automatically expire or delete data; it is used for event-driven processing or replication, not for scheduled data removal. Option C (Global Tables) is wrong because it provides multi-region, fully replicated tables for low-latency access and disaster recovery, but it has no built-in mechanism to expire or delete items based on time. Option D (Point-in-Time Recovery) is wrong because it enables continuous backups of DynamoDB table data to restore to any point within the last 35 days, but it does not delete or manage the lifecycle of individual items.

926
Multi-Selectmedium

Which TWO actions can reduce the cost of an Amazon S3 bucket that stores infrequently accessed data? (Choose 2.)

Select 2 answers
A.Enable cross-region replication
B.Enable versioning to keep multiple versions
C.Use lifecycle policies to expire objects after a certain period
D.Enable MFA Delete for extra security
E.Transition objects to S3 Standard-IA after 30 days
AnswersC, E

Expiration deletes unneeded objects.

Why this answer

Lifecycle policies allow you to define rules that automatically expire (delete) objects after a specified period, which directly reduces storage costs by removing data that is no longer needed. For infrequently accessed data, deleting obsolete objects prevents paying for unnecessary storage over time.

Exam trap

The DEA-C01 exam often tests the misconception that enabling versioning or replication reduces costs, when in fact both increase storage and transfer costs, while lifecycle policies and storage class transitions are the correct cost-saving mechanisms.

927
MCQhard

A data engineer is using AWS Glue Studio to create an ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The source data is in JSON format and contains nested structures. The engineer needs to flatten the nested structures and write the output in Parquet format. The job must be efficient and scalable. Which transformation should the engineer use in Glue Studio to flatten the nested data?

A.Use the 'Relationalize' transform.
B.Use the 'ApplyMapping' transform.
C.Use the 'Filter' transform.
D.Use the 'Join' transform.
AnswerA

The Relationalize transform in AWS Glue flattens nested JSON structures into a relational schema. It unnestes arrays and structs, producing multiple tables that can be joined. This is specifically designed for flattening nested data. It is efficient and scalable because it leverages Spark. The transform outputs a DynamicFrame collection, which can then be written to Parquet. This is the correct choice for flattening nested JSON.

Why this answer

The Relationalize transform in AWS Glue is specifically designed to flatten nested JSON data. It unnests arrays and structs into a relational schema, producing multiple tables. This is the correct transform for the scenario.

ApplyMapping, Filter, and Join do not flatten nested structures. Relationalize is efficient and scalable for this purpose.

Exam trap

The trap here is confusing ApplyMapping with flattening; ApplyMapping only maps fields and does not unnest arrays or structs.

928
Multi-Selecteasy

A data engineer is setting up Amazon S3 event notifications to trigger an AWS Lambda function when new objects are uploaded. Which TWO actions are required to enable this?

Select 2 answers
A.Add a resource-based policy to the Lambda function to allow S3 to invoke it.
B.Enable S3 versioning on the bucket.
C.Create an S3 bucket policy that grants S3 permission to invoke Lambda.
D.Configure an event notification on the S3 bucket for s3:ObjectCreated:* events.
E.Set up an Amazon CloudWatch Events rule to detect S3 uploads.
AnswersA, D

S3 invokes Lambda through a resource-based policy, so the function must grant s3.amazonaws.com permission via AddPermission, scoped to the bucket's source ARN. This satisfies the stem's requirement that S3 be authorised to trigger the function, without which the notification configuration fails with an access error.

Why this answer

Lambda functions use a resource-based policy (also known as a function policy) to grant permissions to other AWS services, such as S3, to invoke the function. Without this policy, S3 will receive an access denied error when trying to trigger the Lambda function. Option D is correct because you must configure an S3 event notification on the bucket for the `s3:ObjectCreated:*` event type to instruct S3 to send a notification to the Lambda function when new objects are uploaded.

Exam trap

The trap here is that candidates often think an S3 bucket policy is needed to allow S3 to invoke Lambda, but in reality, the permission must be granted on the Lambda function's resource-based policy, not on the bucket.

929
MCQmedium

A data engineer is troubleshooting a failed AWS Glue ETL job that reads from an S3 bucket and writes to an Amazon Redshift table. The job fails with a permission error. Which IAM policy addition is MOST likely required for the Glue job's role?

A.Add redshift:DataAPI
B.Add redshift:ModifyCluster
C.Add redshift:DescribeStatement
D.Add redshift:GetClusterCredentials
AnswerD

When Glue writes to Redshift using the Redshift JDBC connector with IAM authentication, the job role must call redshift:GetClusterCredentials to obtain temporary database credentials. Without it, the connection fails with a permission error before any data is written.

Why this answer

AWS Glue jobs that write to Amazon Redshift using the native Redshift connection (not the Redshift Data API) authenticate by calling redshift:GetClusterCredentials to obtain temporary database credentials for the job role. Without that permission, the Glue connection cannot obtain a username/password and the job fails with an access-denied error even though the S3 read side works.

Exam trap

DEA-C01 often tests the confusion between the Redshift JDBC connection path (which needs redshift:GetClusterCredentials) and the Redshift Data API path (which needs redshift-data:* actions), so candidates pick DataAPI permissions for a JDBC-based Glue job.

How to eliminate wrong answers

Option A is wrong because redshift:DataAPI is only needed when using the Redshift Data API (ExecuteStatement) rather than the JDBC connection Glue uses by default. Option B is wrong because redshift:ModifyCluster is an administrative action for changing cluster configuration and has nothing to do with reading or writing table data. Option C is wrong because redshift:DescribeStatement is a Data API action for polling statement status, not for obtaining credentials or connecting via JDBC.

930
MCQmedium

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate a 4 TB on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must complete in a single maintenance window, and the source database cannot be taken offline for more than 30 minutes. The engineer configures a full load plus change data capture (CDC) task. During testing, the full load phase takes 14 hours. Which configuration change will most effectively reduce the time required for the full load phase?

A.Enable parallel load by increasing the number of tables loaded concurrently in the task settings.
B.Increase the Aurora PostgreSQL instance class to the largest available memory-optimized instance.
C.Enable Multi-AZ on the Aurora PostgreSQL cluster before starting the migration task.
D.Change the migration method from full load plus CDC to CDC only.
AnswerA

AWS DMS supports parallel full load, which loads multiple tables simultaneously instead of serially. For a 4 TB Oracle database with many tables, increasing the number of tables loaded concurrently in the task settings significantly reduces overall full load time by using multiple threads and network connections. This is the most effective and supported tuning mechanism for reducing full load duration without changing the source or target.

Why this answer

Parallel full load is the primary DMS tuning lever for reducing bulk migration time. By loading multiple tables concurrently, DMS saturates available network and I/O resources rather than processing tables serially. Switching to CDC only would skip existing data entirely, and scaling the target or enabling Multi-AZ does not address the serial loading bottleneck that dominates a large full load.

Exam trap

The trap here is assuming that scaling the target database or enabling high availability will speed up a DMS full load, when the real bottleneck is the serial table-by-table loading that parallel load addresses.

931
Multi-Selecteasy

A data engineer needs to ingest data from a SaaS application (Salesforce) into Amazon S3 on a daily basis. Which TWO AWS services can be used for this purpose? (Choose TWO.)

Select 2 answers
A.AWS DataSync
B.Amazon Kinesis Data Streams
C.AWS Transfer Family
D.AWS Glue
E.Amazon AppFlow
AnswersD, E

AWS Glue provides built-in Salesforce connectors and can run on a daily schedule, extracting records and writing them directly to Amazon S3. This satisfies the daily ingestion requirement without managing servers, complementing AppFlow as the second valid service.

Why this answer

Amazon AppFlow (option E) is purpose-built for ingesting data from SaaS applications like Salesforce into AWS destinations such as Amazon S3, supporting scheduled daily flows without custom code. AWS Glue (option D) can also perform this ingestion by using its Salesforce connector within Glue ETL jobs or Glue Studio, writing the extracted data to S3 on a schedule. AWS DataSync (A) is designed for migrating and syncing file data between on-premises storage and AWS, not for SaaS API ingestion.

Amazon Kinesis Data Streams (B) is a real-time streaming service that requires custom producers and does not natively connect to Salesforce. AWS Transfer Family (C) only provides managed SFTP, FTPS, and FTP endpoints into S3 or EFS, which does not address SaaS application ingestion.

Exam trap

The trap here is that candidates may confuse AWS DataSync's ability to sync from cloud sources (like EFS) with SaaS applications, but DataSync does not support Salesforce or other SaaS APIs natively.

932
MCQmedium

A data engineer is managing an Amazon Redshift cluster and needs to load data from Amazon S3. The engineer uses the COPY command but encounters the error: 'S3ServiceException: Access Denied.' The Redshift cluster has an IAM role attached with permissions to access the S3 bucket. What is the most likely cause of this error?

A.The IAM role attached to the Redshift cluster does not have the correct trust relationship with Redshift.
B.The Redshift cluster is not in the same AWS Region as the S3 bucket.
C.The COPY command is missing the CREDENTIALS parameter.
D.The S3 bucket policy does not grant access to the IAM role used by Redshift.
AnswerD

Even if the IAM role has S3 permissions, the S3 bucket policy must also allow access from that role. If the bucket policy explicitly denies or does not allow the role, Redshift will receive an Access Denied error. This is a common misconfiguration when the bucket is in a different account or has restrictive policies.

Why this answer

The Access Denied error from S3 during a COPY operation typically indicates that the IAM role used by Redshift does not have sufficient permissions, or the S3 bucket policy does not allow the role. Even if the role has an S3 policy, an explicit deny or missing allow in the bucket policy can cause this. The engineer should verify both the IAM role's policies and the S3 bucket policy.

Exam trap

The trap here is focusing only on the IAM role's permissions and ignoring the S3 bucket policy, which can also deny access.

933
MCQeasy

A data engineer must choose a storage service for a new application that requires single-digit millisecond latency at any scale, a flexible schema, and automatic scaling of throughput without provisioning capacity. The access pattern is key-value lookups by user ID with occasional range queries on a sort key. Which AWS service should the engineer select?

A.Amazon RDS for PostgreSQL with automatic storage scaling
B.Amazon S3 with S3 Select
C.Amazon Redshift with a distribution key on user ID
D.Amazon DynamoDB with on-demand capacity mode
AnswerD

DynamoDB delivers consistent single-digit millisecond latency at any scale, supports a flexible schema with a partition key and optional sort key, and on-demand capacity mode scales read and write throughput automatically without provisioning. Key-value lookups by user ID map to the partition key, and range queries map to the sort key, matching the access pattern exactly.

Why this answer

DynamoDB is purpose-built for key-value and document workloads needing consistent single-digit millisecond latency at any scale. A partition key on user ID serves point lookups, and a sort key supports efficient range queries. On-demand capacity mode removes provisioning and scales automatically with traffic, while the flexible schema accommodates evolving attributes without migrations, matching every stated requirement.

Exam trap

The trap here is assuming a relational database such as Amazon RDS or a warehouse such as Redshift can deliver single-digit millisecond latency at any scale, when that guarantee belongs to DynamoDB's key-value design.

934
MCQhard

A healthcare company stores patient records in an S3 bucket encrypted with SSE-S3. The data engineering team uses AWS Glue ETL jobs to process this data and load it into an Amazon Redshift cluster for analytics. Recently, the security team mandated that all sensitive data must be encrypted at rest using customer-managed keys (CMK) in AWS KMS, and that the keys must be rotated automatically every year. The team updated the S3 bucket to use SSE-KMS with a CMK and enabled automatic key rotation. However, after the change, the Glue ETL jobs that read from the S3 bucket started failing with 'Access Denied' errors. The Glue job uses an IAM role named 'GlueETLRole' that has the following permissions: s3:GetObject on the bucket, kms:Decrypt and kms:GenerateDataKey on the CMK, and all necessary Glue permissions. The Redshift cluster is also encrypted with a different CMK, and the Glue role has kms:Decrypt on that key as well. What is the most likely cause of the failure?

A.The KMS key policy for the CMK used for S3 encryption does not grant 'GlueETLRole' permission to use the key.
B.The IAM role 'GlueETLRole' does not have kms:Decrypt permission on the CMK used for S3 encryption.
C.The Glue job requires kms:Encrypt permission to read encrypted data from S3.
D.The S3 VPC endpoint policy does not allow the Glue job to access the KMS key.
AnswerA

The CMK's key policy must explicitly allow the GlueETLRole to call kms:Decrypt and kms:GenerateDataKey; IAM permissions alone are insufficient because KMS authorises against both the key policy and the identity policy. Since SSE-KMS now wraps every S3 object read, the missing key policy grant produces the Access Denied failures.

Why this answer

The most likely cause is that the KMS key policy for the CMK used for S3 encryption does not grant 'GlueETLRole' permission to use the key. Even if the IAM role has kms:Decrypt and kms:GenerateDataKey permissions, the KMS key policy must also allow the role to use the key. KMS requires both IAM policy and key policy to grant access.

Exam trap

DEA-C01 often tests the dual-authorization requirement of KMS, and candidates may incorrectly assume that IAM permissions alone are sufficient, overlooking the need for key policy grants.

How to eliminate wrong answers

Option B is wrong because the scenario states the IAM role already has kms:Decrypt and kms:GenerateDataKey on the CMK, so the IAM policy is not the issue. Option C is wrong because reading encrypted data from S3 requires kms:Decrypt, not kms:Encrypt; kms:Encrypt is for writing encrypted data. Option D is wrong because S3 VPC endpoint policies control network access to S3, not access to KMS keys; KMS access is governed by IAM and key policies.

935
MCQhard

A data engineer manages an Amazon Redshift cluster that experiences performance degradation during complex analytical queries. The engineer notices that some queries spill to disk. The engineer wants to improve query performance by optimizing the distribution style and sort keys. Which action should the engineer take first?

A.Change all tables to use EVEN distribution style to balance data across nodes.
B.Analyze the query execution plans to identify tables with high redistribution or broadcast steps.
C.Set the sort key to the column most frequently used in WHERE clauses for all tables.
D.Increase the number of nodes in the cluster to add more memory and storage.
AnswerB

Query execution plans reveal how data is distributed and joined. High redistribution or broadcast steps indicate suboptimal distribution styles, causing data movement across nodes. By analyzing these plans, the engineer can identify which tables need distribution key changes. This is the first step before making changes, as it provides data-driven insights. It ensures that modifications target the actual bottlenecks.

Why this answer

Analyzing query execution plans is the essential first step to identify performance bottlenecks such as data redistribution or broadcast operations. These indicate suboptimal distribution styles. Once identified, the engineer can make informed changes to distribution keys and sort keys.

The other options are either premature or not targeted at the root cause.

Exam trap

The trap here is jumping to schema changes or scaling without first diagnosing the actual query execution bottlenecks.

936
MCQmedium

A company needs to automate the detection of sensitive data in Amazon S3 and generate reports. Which AWS service should be used?

A.Amazon Macie
B.Amazon Inspector
C.Amazon GuardDuty
D.AWS Config
AnswerA

Amazon Macie uses machine learning and pattern matching to discover sensitive data such as personally identifiable information across S3 buckets, satisfying the automated detection requirement. It continuously evaluates bucket inventory, assigns severity-based findings, and produces reports, directly meeting the reporting constraint without custom code or manual review.

Why this answer

Amazon Macie is a fully managed data security and data privacy service that uses machine learning and pattern matching to discover and protect sensitive data in Amazon S3. It automatically detects personally identifiable information (PII), financial data, and credentials, and generates detailed findings and reports. This directly matches the requirement to automate sensitive data detection and reporting in S3.

Exam trap

DEA-C01 often tests the confusion between Macie (data discovery in S3), Inspector (vulnerability scanning), and GuardDuty (threat detection), tempting candidates to pick Inspector for 'sensitive data' because it sounds security-related.

How to eliminate wrong answers

Option B is wrong because Amazon Inspector is a vulnerability management service that scans EC2 instances, container images, and Lambda functions for software vulnerabilities and network exposure — it does not detect sensitive data in S3. Option C is wrong because Amazon GuardDuty is a threat detection service that monitors for malicious activity and unauthorized behavior using VPC Flow Logs, CloudTrail, and DNS logs — not data classification. Option D is wrong because AWS Config is a configuration compliance service that records resource changes and evaluates them against rules, but it does not perform data-level sensitive content discovery.

937
MCQhard

A data engineer maintains an AWS Glue Data Catalog with databases for several lines of business. Auditors require that every change to table definitions, partition additions, and schema edits in the catalog be recorded with the identity of the caller and that the records be retained for 365 days in a dedicated S3 bucket. Which solution should the data engineer implement?

A.Enable AWS Glue job metrics and continuous logging, then export the logs to Amazon CloudWatch Logs with a retention period of 365 days.
B.Configure an AWS CloudTrail trail that logs management events, including AWS Glue Data Catalog API calls, and deliver the trail to the dedicated S3 bucket with a 365-day lifecycle rule.
C.Turn on AWS Config recording for the AWS::Glue::Table resource type and store configuration snapshots in the dedicated S3 bucket for 365 days.
D.Enable AWS CloudTrail data events for the AWS Glue Data Catalog and deliver the logs to the dedicated S3 bucket with a 365-day lifecycle retention policy.
AnswerB

Data Catalog operations such as CreateTable, UpdateTable, and BatchCreatePartition are management events recorded by CloudTrail. A trail that logs management events captures the caller identity and request details, and delivering to a dedicated S3 bucket with a 365-day lifecycle rule meets the retention requirement for auditors.

Why this answer

AWS Glue Data Catalog mutations such as CreateTable, UpdateTable, and BatchCreatePartition are management events, and a CloudTrail trail that logs management events records the caller identity and request details for each. Delivering that trail to a dedicated S3 bucket with a 365-day lifecycle rule satisfies both the identity-capture and retention requirements, whereas data events, job logs, and AWS Config do not cover these catalog operations.

Exam trap

The trap here is confusing CloudTrail data events with management events, when Data Catalog changes are management events and are not captured by enabling data events.

938
MCQhard

Refer to the exhibit. A data engineer has configured an S3 event notification to send an event to an SQS queue when objects are created in the 'incoming/' prefix. The engineer wants to trigger an AWS Lambda function to process the object. However, the Lambda function is not being invoked. What is the most likely cause?

A.The Lambda function lacks permission to read from the S3 bucket.
B.Lambda is not configured as an event source for the SQS queue.
C.The SQS queue does not exist or is in a different account.
D.The S3 event notification filter prefix is incorrect.
AnswerB

An SQS queue only buffers messages; Lambda must be configured with the queue as an event source mapping to poll and invoke. Without that mapping, notifications accumulate unprocessed, so the function never runs despite S3 delivering events correctly.

Why this answer

The Lambda function is not being invoked because the SQS queue is not configured as an event source for Lambda. Even though S3 sends events to SQS, Lambda will only poll and process messages from the queue if an event source mapping (e.g., via CreateEventSourceMapping) is established. Without this mapping, the messages sit in the queue and Lambda remains idle.

Exam trap

The DEA-C01 exam often tests the distinction between sending events to a queue (S3 → SQS) and actually consuming them (SQS → Lambda), so candidates mistakenly assume that simply having S3 send to SQS will automatically trigger Lambda without an explicit event source mapping.

How to eliminate wrong answers

Option A is wrong because the Lambda function does not need permission to read from the S3 bucket directly; it only needs permission to read from the SQS queue (via the event source mapping) and to access the object in S3 once triggered. Option C is wrong because if the SQS queue did not exist or was in a different account, the S3 event notification would fail immediately (S3 would log an error), but the question states the event is sent to SQS, implying the queue exists and is accessible. Option D is wrong because the S3 event notification filter prefix 'incoming/' is correctly configured to match objects created under that prefix; if it were incorrect, no events would be sent to SQS at all.

939
MCQeasy

A data engineer needs to back up an Amazon DynamoDB table daily. The backup must be restorable to a specific point in time within the last 24 hours. Which solution meets these requirements with the LEAST operational overhead?

A.Create an on-demand backup of the table every 24 hours.
B.Use DynamoDB Streams to replicate data to another table.
C.Enable point-in-time recovery (PITR) on the table.
D.Export the table data to Amazon S3 every 6 hours using a Lambda function.
AnswerC

Point-in-time recovery continuously backs up the table, enabling restore to any second within the last 35 days with no scripts or scheduling. This satisfies the 24-hour restore requirement and minimises operational overhead compared with on-demand backups or custom export pipelines.

Why this answer

DynamoDB point-in-time recovery (PITR) provides continuous backups of your table data with a restore granularity of one second, allowing you to restore to any point within the last 35 days. It requires no manual intervention, thus offering the least operational overhead. Enabling PITR meets the requirement to restore to a specific point in time within the last 24 hours.

Exam trap

DEA-C01 often tests the misconception that DynamoDB Streams or on-demand backups provide point-in-time recovery, when only PITR offers continuous, second-level restore within the retention window.

How to eliminate wrong answers

Option A is wrong because on-demand backups are manual and only capture the state at the time of backup; they do not provide point-in-time recovery within a 24-hour window. Option B is wrong because DynamoDB Streams captures item-level changes but does not provide a built-in restore mechanism; replicating to another table would require custom code and does not offer point-in-time restore. Option D is wrong because exporting to S3 every 6 hours is a batch process that does not provide continuous point-in-time recovery and involves significant operational overhead to manage exports and restores.

940
MCQmedium

Refer to the exhibit. A data engineer runs this AWS CLI command to execute an Athena query. What is the purpose of the EncryptionConfiguration parameter?

A.It encrypts the query string in transit
B.It encrypts the data in the source table
C.It enables client-side encryption for the query output
D.It encrypts the query results stored in Amazon S3 at rest
AnswerD

EncryptionConfiguration specifies SSE-S3 or SSE-KMS settings applied to the query result files Athena writes into the S3 query-result location. This satisfies the parameter's purpose: protecting results at rest in Amazon S3, not encrypting data in transit or the source data being queried.

Why this answer

The EncryptionConfiguration parameter in Athena specifies how the query results stored in S3 are encrypted at rest. SSE_S3 means server-side encryption with S3-managed keys. It does not encrypt the query itself, data in transit, or the source data.

941
MCQhard

A data engineer maintains an AWS Glue ETL job that reads JSON from Amazon S3 and writes Parquet to a second bucket. Downstream consumers report that numeric fields occasionally arrive as strings and timestamps are sometimes null. The engineer must make the job resilient to these schema variations without failing the run. Which approach should the engineer take?

A.Switch the output format to CSV so type coercion is handled by the consumer
B.Use a ResolveChoice transform to cast columns and a FillMissingValues transform to populate null timestamps
C.Enable the AWS Glue Data Catalog schema evolution option and set the job to merge schemas from the source
D.Increase the number of DPU workers and enable job bookmarks
AnswerB

ResolveChoice resolves ambiguous or mixed types such as a column that is sometimes a string and sometimes a number, applying a cast or type promotion so the output Parquet has a consistent schema. FillMissingValues then supplies defaults for null timestamps. Together they let the job complete and emit uniform Parquet instead of failing or producing inconsistent types downstream.

Why this answer

Mixed types and null timestamps are data-quality issues that must be normalized inside the transformation. ResolveChoice reconciles a column that arrives as both string and numeric into a single declared type, and FillMissingValues substitutes defaults for absent timestamps. Together they let the job emit consistent Parquet without failing on schema variance.

Exam trap

The trap here is reaching for catalog-level schema evolution or more compute when the real problem is per-record type inconsistency that only a transform such as ResolveChoice can reconcile.

942
MCQmedium

A data engineer manages an Amazon Redshift cluster that runs a nightly ETL load followed by complex analytical queries. Users report that queries during the day are slower than expected, and the team wants to isolate the ETL workload so it cannot consume resources needed by the analytical queries. The cluster uses provisioned nodes. What is the MOST appropriate solution?

A.Increase the number of nodes in the cluster to provide more resources for all workloads.
B.Enable concurrency scaling on the cluster and route all queries through it.
C.Create a separate Redshift workload management (WLM) queue for the ETL role and assign the analytical queries to a different queue.
D.Move the ETL workload to a separate Redshift cluster and use Amazon Redshift Spectrum for analytics.
AnswerC

Redshift WLM lets you define multiple queues with dedicated memory and concurrency slots, and you can route queries to queues based on user groups or query groups. Assigning ETL to its own queue prevents it from starving the analytics queue. This directly isolates the workloads on the same provisioned cluster without extra infrastructure.

Why this answer

Redshift WLM provides workload isolation by letting you define multiple queues, each with its own memory allocation and query concurrency, and by routing queries to queues based on user groups or query groups. Placing the ETL job in a dedicated queue ensures it cannot consume the memory and slots reserved for analytical queries, which addresses the slowdowns without adding clusters or nodes.

Exam trap

The trap here is assuming that adding nodes or enabling concurrency scaling will isolate workloads, when isolation requires separate WLM queues with defined memory and concurrency.

943
MCQmedium

A media company stores video files in an S3 bucket. The files are processed by a fleet of EC2 instances that read the files, add watermarks, and write the output back to the same bucket. Recently, the processing jobs have been failing with '500 Internal Server Error' and '503 Slow Down' errors. The data engineer checks the S3 bucket metrics and sees that the PUT/GET request rate is consistently above 5,500 requests per second for a single prefix. The engineer needs to resolve the errors with minimal changes to the application code. Which course of action should the engineer take?

A.Use S3 Batch Operations to process the files.
B.Increase the number of EC2 instances to process files in parallel.
C.Enable S3 Transfer Acceleration on the bucket to improve throughput.
D.Modify the application to add a random hash prefix to the object keys to distribute load across multiple prefixes.
AnswerD

S3 partitions request throughput per prefix, so exceeding roughly 3,500 PUT or 5,500 GET requests per second on one prefix triggers 503 Slow Down throttling. Adding a random hash prefix spreads keys across many partitions, restoring throughput without changing the bucket or application logic significantly.

Why this answer

S3 automatically partitions request traffic by key prefix, and each partition supports roughly 3,500 PUT/COPY/POST/DELETE or 5,500 GET/HEAD requests per second. When a single prefix exceeds these limits, S3 returns 503 Slow Down errors. Adding a random hash prefix to object keys distributes the load across many prefixes, each with its own request rate limit, thereby eliminating throttling without changing the core application logic.

This is the recommended best practice for high-throughput workloads.

Exam trap

DEA-C01 often tests the misconception that adding more compute resources (EC2 instances) or enabling Transfer Acceleration will solve S3 throttling, when the actual fix is to distribute the request load across multiple key prefixes.

How to eliminate wrong answers

Option A is wrong because S3 Batch Operations is designed for bulk operations on existing objects (like copying or tagging) and does not increase request rate limits or resolve throttling for ongoing PUT/GET operations. Option B is wrong because adding more EC2 instances would increase the request rate further, exacerbating the throttling issue rather than alleviating it. Option C is wrong because S3 Transfer Acceleration speeds up transfers over long distances by using AWS edge locations, but it does not increase the per-prefix request rate limits and will not resolve 503 Slow Down errors caused by exceeding those limits.

944
MCQeasy

A data engineer is troubleshooting an AWS Glue job that reads from an Apache Kafka topic using a Glue connector. The job fails with 'TimeoutException'. The Kafka cluster is in a VPC. Which step should the engineer take FIRST?

A.Check the security group and network ACLs associated with the Glue job's VPC.
B.Increase the Kafka consumer session timeout.
C.Update the Glue connector to the latest version.
D.Change the Glue job type from Spark to Python Shell.
AnswerA

A TimeoutException when a Glue connector reaches a Kafka cluster inside a VPC most often means connectivity is blocked, not that Kafka is misconfigured. Security groups and network ACLs are the first hop to verify, since either can silently drop traffic on the required broker ports.

Why this answer

A TimeoutException from a Glue job reading Kafka in a VPC almost always indicates a network connectivity problem between the Glue job's ENIs and the Kafka brokers. The first step is to verify that the security groups and network ACLs allow traffic on the Kafka ports (typically 9092/9093) between the Glue job's subnet and the Kafka cluster's subnet. This is the most common root cause and the fastest to check.

Exam trap

The trap is jumping to application-layer fixes (session timeout, connector version, job type) when the error is a network-layer TimeoutException — candidates forget that in AWS, 'timeout' almost always means security group, NACL, or route table, not a code or configuration parameter.

How to eliminate wrong answers

Option B is wrong because increasing the Kafka consumer session timeout only helps if the consumer is being evicted due to slow processing — it does not fix a network-level timeout where packets never reach the broker, and it would mask rather than solve the problem. Option C is wrong because updating the Glue connector version is a change-management action that does not address the underlying connectivity failure; connector version issues typically manifest as class-not-found or deserialization errors, not TimeoutException. Option D is wrong because changing the job type from Spark to Python Shell does not change the network path — the Python Shell environment still needs VPC connectivity to reach Kafka, so the timeout would persist.

945
MCQmedium

A data engineer notices that an AWS Glue ETL job that processes streaming data from Amazon Kinesis Data Streams is failing intermittently with a 'ResourceNotFoundException' error for the Kinesis stream. The job has been running successfully for weeks. Which action should the engineer take to resolve the issue?

A.Increase the number of shards in the Kinesis data stream to handle higher throughput.
B.Rename the Kinesis data stream to match the stream name used in the Glue job exactly, including case.
C.Add the 'kinesis:DescribeStream' permission to the IAM role used by the Glue job.
D.Increase the timeout for the Glue job in the job configuration.
AnswerC

Missing DescribeStream permission causes intermittent resource not found errors.

Why this answer

A 'ResourceNotFoundException' from a Glue job reading Kinesis typically indicates the Glue execution role lacks permission to describe or access the stream, causing the service to report the stream as not found rather than as access-denied. Adding 'kinesis:DescribeStream' (and typically DescribeStreamSummary, GetShardIterator, GetRecords) to the IAM role resolves the issue. Since the job ran successfully for weeks, a recent IAM policy change or role modification is the likely trigger.

Exam trap

The trap is assuming 'ResourceNotFoundException' always means the resource is missing or misnamed — in AWS, permission denials are frequently masked as not-found errors, so candidates who jump to renaming or recreating resources miss the IAM root cause.

How to eliminate wrong answers

Option A is wrong because increasing shards addresses throughput scaling, not a ResourceNotFoundException — more shards would not fix a permissions or naming issue and could actually complicate the problem. Option B is wrong because if the stream name were mismatched, the job would have failed from day one, not after weeks of successful runs; the question explicitly states the job has been running successfully. Option D is wrong because a timeout produces a different error (e.g., 'JobRunTimedOut' or a timeout exception), not ResourceNotFoundException, and timeouts don't manifest as missing-resource errors.

946
MCQhard

A data engineer is using AWS Glue Studio to build a job that reads from an Amazon Kinesis Data Stream, performs a 5-minute tumbling window aggregation, and writes results to Amazon S3. The job must run continuously and handle late-arriving records within the window. Which configuration should the engineer use?

A.Configure the job with the Glue streaming ETL type, set the window size to 5 minutes, and enable checkpointing to track stream position
B.Use a Glue batch job with the 'groupBy' transform set to a 5-minute window and run it once per hour
C.Use an AWS Lambda function triggered by the stream with a 5-minute timeout to aggregate records and write to S3
D.Configure the job as a batch job with a schedule trigger that runs every 5 minutes and reads the stream's latest records
AnswerA

Glue streaming ETL jobs run continuously, consume from Kinesis, and support windowed aggregations with a configurable window size. Checkpointing tracks the stream position so the job resumes correctly and handles late records within the window as designed.

Why this answer

Glue streaming ETL is purpose-built for continuous consumption from Kinesis and supports windowed aggregations with checkpointing. Selecting the streaming job type with a 5-minute window and enabled checkpoints gives the continuous, late-tolerant behavior the scenario requires, unlike scheduled batch jobs or Lambda.

Exam trap

The trap here is choosing a scheduled batch job because the window is 5 minutes, when only a Glue streaming ETL job provides continuous windowed aggregation with checkpointing.

947
MCQeasy

A data engineer needs to run a Python-based transformation on each object as it lands in an Amazon S3 bucket. The objects are small (under 10 MB), arrive sporadically, and must be processed within seconds. Which approach is MOST appropriate?

A.Schedule an AWS Glue crawler and ETL job to run every 15 minutes over the bucket.
B.Run an Amazon EMR cluster with a step function that polls the bucket every minute.
C.Use Amazon Kinesis Data Firehose to deliver objects to a processing endpoint.
D.Configure an S3 event notification to invoke an AWS Lambda function that performs the transformation.
AnswerD

S3 event notifications can trigger Lambda directly on object creation, giving near-real-time, per-object processing without managing servers. Lambda handles small payloads well within its 15-minute limit, and you pay only per invocation, which matches sporadic arrivals and sub-10 MB objects perfectly.

Why this answer

S3 event notifications triggering Lambda is the standard serverless pattern for reacting to individual object uploads in near real time. It avoids polling, scales automatically, and suits small files and sporadic arrivals. Glue, Firehose, and EMR introduce batching, buffering, or cluster overhead that conflicts with the seconds-level, per-object requirement.

Exam trap

The trap here is reaching for a managed ETL service like Glue when the requirement is event-driven per-object processing with seconds-level latency.

948
Multi-Selecteasy

A company is using AWS Glue to process data stored in Amazon S3. The Glue job runs successfully but takes longer than expected. Which TWO actions can reduce the job runtime?

Select 2 answers
A.Disable job bookmarking
B.Increase the number of DPUs allocated to the job
C.Reduce the number of workers
D.Change the job type from Spark to Python shell
E.Partition the input data in S3
AnswersB, E

Glue allocates DPUs per job, and each DPU supplies processing capacity and memory. Raising the DPU count lets more executors run partitions concurrently, reducing runtime for large shuffles or skewed workloads, provided the data is partitioned into enough input splits.

Why this answer

Option B is correct because increasing the number of DPUs (Data Processing Units) allocated to a Glue job adds more compute capacity and parallelism, allowing Spark executors to process partitions concurrently and thereby reducing overall runtime. Option E is correct because partitioning the input data in S3 (for example, by date or category) lets Glue's Spark engine prune irrelevant partitions and read only the data it needs, cutting I/O and shuffle overhead. Option A is wrong because disabling job bookmarks causes Glue to reprocess already-processed data, which typically increases runtime rather than reducing it.

Option C is wrong because reducing the number of workers lowers parallelism and generally makes the job slower. Option D is wrong because Python shell jobs are single-node, non-distributed, and intended for lightweight scripts, so they cannot efficiently process large datasets that a distributed Spark job handles.

949
MCQhard

A data engineer runs a weekly AWS Glue ETL job that processes data from Amazon DynamoDB to Amazon S3. The job reads the entire table every time, which is slow and expensive. The job needs to process only items that changed since the last run. Which solution should the engineer implement?

A.Use DynamoDB Scan with a LastEvaluatedKey to paginate and store the last scanned key to resume next time
B.Enable DynamoDB Streams and process change events with AWS Lambda to write to S3
C.Add a Global Secondary Index (GSI) on a timestamp attribute and query only new records
D.Use AWS Database Migration Service (DMS) with ongoing replication from DynamoDB to S3
AnswerB

DynamoDB Streams captures item-level changes in near real time, and Lambda consumes those events to write only modified items to Amazon S3. This satisfies the requirement to process only items changed since the last run, eliminating the slow, expensive full-table scans the weekly Glue job currently performs.

Why this answer

DynamoDB Streams captures item-level changes (inserts, updates, deletes) in near real-time. An AWS Lambda function can process these change events and write only the incremental data to Amazon S3, eliminating the need to scan the entire DynamoDB table. This approach is both cost-effective and efficient for incremental data ingestion.

Exam trap

The trap here is that candidates may think a GSI on a timestamp (Option C) is sufficient for incremental processing, but it fails to capture updates to existing items that do not change the timestamp, and it still requires a full scan of the index to find new records.

How to eliminate wrong answers

Option A is wrong because using DynamoDB Scan with LastEvaluatedKey still reads the entire table every time; it only paginates the results, not reduces the data read. Option C is wrong because a Global Secondary Index (GSI) on a timestamp attribute does not automatically track changes; you would still need to query for new records based on a timestamp, which requires storing the last processed timestamp and can miss updates to existing items. Option D is wrong because AWS Database Migration Service (DMS) with ongoing replication is designed for continuous database migration and can be complex to set up for simple incremental loads; it is overkill compared to using DynamoDB Streams with Lambda, and DMS does not natively write to S3 in a format optimized for analytics without additional transformation.

950
MCQmedium

A data engineer is building a data lake on Amazon S3 and must enforce that all objects containing personally identifiable information are encrypted with a customer managed AWS KMS key, while allowing automatic key rotation and audit of key usage. Objects must remain readable by an AWS Glue job and an Amazon Athena workgroup. Which configuration should the engineer choose?

A.Client-side encryption with an application managed key stored in the application code
B.SSE-KMS with the AWS managed key aws/s3 and default bucket encryption
C.SSE-KMS with a customer managed key, key rotation enabled, and bucket policy enforcing the key
D.SSE-S3 with bucket versioning enabled
AnswerC

SSE-KMS with a customer managed key gives the organization control over the key policy, enables automatic annual rotation, and records every cryptographic operation in AWS CloudTrail for audit. A bucket policy that denies uploads not using the specified key enforces consistency. Granting the Glue role and Athena workgroup kms:Decrypt and kms:GenerateDataKey permissions allows both services to read and write objects.

Why this answer

A customer managed KMS key with SSE-KMS satisfies control, rotation, and auditability: the key policy scopes who may use it, automatic rotation can be enabled, and CloudTrail logs every Encrypt, Decrypt, and GenerateDataKey call. A bucket policy that denies noncompliant uploads enforces the standard, and granting the Glue role and Athena workgroup decrypt permissions keeps analytics working. AWS managed keys and SSE-S3 cannot meet the customer control and rotation requirements.

Exam trap

The trap here is treating any KMS-based encryption as equivalent, when only a customer managed key provides configurable rotation and a key policy that can be audited and restricted.

951
MCQeasy

A data engineer needs to encrypt data in transit between an Amazon RDS for MySQL instance and an application. Which solution should be used?

A.Enable encryption at rest using AWS KMS
B.Use SSL/TLS to connect to the RDS instance
C.Store the data in Amazon S3 with server-side encryption
D.Use AWS CloudHSM to generate and store encryption keys
AnswerB

SSL/TLS encrypts the wire protocol between the application and the RDS for MySQL endpoint, satisfying the in-transit encryption requirement. RDS MySQL supports TLS connections natively; clients negotiate certificates during the handshake, protecting credentials and query data from interception. Encryption at rest options such as KMS keys do not address data moving across the network.

Why this answer

Encryption in transit for an RDS for MySQL instance is achieved by connecting with SSL/TLS, which encrypts the wire protocol between the application and the database. RDS MySQL supports TLS and provides an rds-ca certificate that the client uses to verify the server. Enabling SSL/TLS on the connection is the direct, correct answer for data in transit.

Exam trap

DEA-C01 often tests whether candidates confuse encryption at rest (KMS, SSE) with encryption in transit (TLS/SSL) — the phrase 'in transit' is the key discriminator, and options mentioning KMS or S3 encryption are the classic distractors.

How to eliminate wrong answers

Option A is wrong because AWS KMS encryption at rest protects data on the underlying storage volume, not the network path between the application and the database. Option C is wrong because storing data in S3 with SSE is a different storage service and does not encrypt the RDS connection; it also changes the architecture rather than securing the existing path. Option D is wrong because CloudHSM generates and stores keys for encryption at rest or custom key operations, not for encrypting the MySQL client-server transport.

952
MCQmedium

A data engineer needs to transform JSON data into Parquet format using AWS Glue. The input data has nested fields. Which Glue feature should be used to flatten the nested structure?

A.Relationalize transform
B.DropNullFields transform
C.FindMatches transform
D.Map transform
AnswerA

Relationalize is a Glue transform that converts nested, semi-structured data into a set of flat relational tables, unnesting arrays and structs into separate tables linked by join keys. It directly addresses flattening nested JSON before writing Parquet.

Why this answer

The Relationalize transform is the correct choice because it is specifically designed to flatten nested JSON structures (such as arrays and structs) into a set of related tables (DataFrames) that can be written as Parquet. This transform recursively extracts nested fields, creating separate DataFrames for each level of nesting, which is essential for converting complex JSON into a flat, columnar Parquet format suitable for analytics.

Exam trap

The trap here is that candidates may confuse the Map transform (which can flatten JSON with custom code) with a built-in flattening feature, but AWS Glue's Relationalize is the dedicated, no-code solution for this specific task, and the exam expects you to know the exact purpose of each transform.

How to eliminate wrong answers

Option B (DropNullFields transform) is wrong because it only removes fields with null values from the schema, not flattening nested structures. Option C (FindMatches transform) is wrong because it is used for fuzzy matching and deduplication of records, not for schema transformation. Option D (Map transform) is wrong because it applies a custom function to each row or column but does not inherently flatten nested JSON; it requires manual coding to handle nesting, whereas Relationalize automates the process.

953
MCQmedium

A data engineering team uses Amazon S3 to store raw data files. They have an AWS Glue ETL job that reads from an S3 bucket, transforms the data, and writes to a Redshift cluster. The job runs daily and has been failing intermittently with the error: 'An error occurred while calling o143.pyWriteDynamicFrame. S3 Access Denied'. The team has confirmed that the IAM role used by the Glue job has s3:GetObject and s3:PutObject permissions on the bucket and all objects. The Redshift cluster is in the same VPC and the Glue connection is configured correctly. What is the most likely cause of the failure?

A.The Redshift cluster is not publicly accessible and the Glue job does not have a VPC endpoint to Redshift.
B.The Glue job has exceeded the maximum execution time and is being killed by AWS.
C.The Glue job is using the wrong JDBC driver version for Redshift.
D.The Glue job's IAM role lacks permission to write to the Glue temporary file bucket (aws-glue-*).
AnswerD

Glue writes intermediate results to its own temporary S3 bucket before loading into Redshift, so the pyWriteDynamicFrame Access Denied arises there, not on the source or target. The role's GetObject and PutObject grants on the data bucket do not cover the aws-glue-* temporary bucket.

Why this answer

AWS Glue writes intermediate data to a temporary S3 bucket (aws-glue-* in the same region) before loading into Redshift. Even though the job has S3 permissions on the source bucket, the IAM role must also have s3:GetObject, s3:PutObject, and s3:DeleteObject on the Glue temporary bucket. The 'S3 Access Denied' error during pyWriteDynamicFrame indicates the write to this temp bucket is failing.

Exam trap

The trap is focusing on the source bucket permissions and Redshift connectivity while overlooking that Glue requires separate permissions for its temporary staging bucket, which is a frequent cause of 'S3 Access Denied' errors.

How to eliminate wrong answers

Option A is wrong because the error is an S3 access denied, not a Redshift connectivity issue; the Glue connection is confirmed correct and the cluster is in the same VPC. Option B is wrong because exceeding execution time would produce a timeout or job cancellation error, not an S3 Access Denied. Option C is wrong because a wrong JDBC driver would cause a connection or SQL error, not an S3 permission error.

954
MCQmedium

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate a large on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must minimize downtime, so the engineer needs to capture ongoing changes while the initial full load runs. Which AWS DMS task configuration should the engineer use?

A.Full load only, followed by a separate full load task after cutover.
B.Change data capture (CDC) only, without an initial full load.
C.Full load plus CDC, but disable the target schema creation.
D.Full load plus change data capture (CDC) in a single task.
AnswerD

A full load plus CDC task first copies existing rows and then continuously applies ongoing changes captured from the source transaction logs, keeping the target in sync until cutover. This minimizes downtime because the target is current at the moment of switchover, which is exactly what the low-downtime requirement demands.

Why this answer

To minimize downtime during a database migration, AWS DMS should run a task that performs a full load and then continues with change data capture. The full load seeds the target with existing rows, and CDC applies ongoing inserts, updates, and deletes from the source logs, keeping Aurora PostgreSQL synchronized until the cutover window, which can then be brief.

Exam trap

The trap here is thinking that a full load alone keeps the target current, when in fact changes made during and after the load require change data capture to be applied.

955
MCQmedium

A data engineer manages an AWS Glue ETL job that writes Parquet files to Amazon S3. Downstream Amazon Athena queries started returning duplicate rows after the job was modified to enable job bookmarks. The job reads from an S3 source prefix where new files are appended hourly and the transformation includes a join that reorders records. Which action will most reliably eliminate the duplicate rows while preserving incremental processing?

A.Change the output format from Parquet to CSV so Athena can deduplicate rows during query execution.
B.Disable job bookmarks and reprocess the entire source prefix on every run.
C.Reset the job bookmark state and verify the transformation produces deterministic output for each input file.
D.Increase the number of Glue DPUs allocated to the job so the join completes in a single task.
AnswerC

Job bookmarks track which S3 objects were already processed, so when a transformation changes the input-to-output relationship, stale bookmark state can cause overlaps. Resetting the bookmark forces a clean reprocessing baseline, and making the transformation deterministic per input file ensures each source object maps to exactly one output. Together these restore correct incremental behavior without reprocessing everything on every run.

Why this answer

Job bookmarks persist state about which S3 objects have been processed. When a transformation changes so that one input object no longer maps cleanly to one output, stale bookmark state causes records to be reprocessed and duplicated. Resetting the bookmark establishes a known baseline, and ensuring the transformation is deterministic per input file keeps each object processed exactly once on subsequent incremental runs.

Exam trap

The trap here is assuming that enabling job bookmarks automatically guarantees exactly-once output regardless of how the transformation is written.

956
MCQeasy

A data engineer needs to schedule an AWS Glue ETL job to run every hour and process new data that arrives in an S3 bucket. The job should only process files that have been added since the last run. Which approach should the engineer use to track which files have been processed?

A.Configure S3 Event Notifications to trigger the Glue job on each new object creation.
B.Enable job bookmarks in the Glue ETL job.
C.Store the last processed timestamp in a DynamoDB table and query it at the start of the job.
D.Use S3 Inventory to list all objects and filter by last modified date in the job.
AnswerB

Job bookmarks persist state across runs, letting Glue track previously processed S3 objects by path and last-modified timestamp. This satisfies the hourly incremental requirement: each run reads only files added since the prior execution, avoiding full reprocessing. No external tracking table or custom filtering logic is needed.

Why this answer

AWS Glue job bookmarks automatically track the last processed data, allowing the job to incrementally process only new files since the last run. This is the native, built-in mechanism for stateful incremental processing in Glue ETL, eliminating the need for external tracking.

Exam trap

The trap here is that candidates often choose Option A (S3 Event Notifications) because it seems like a direct trigger for new files, but they overlook that the requirement is for an hourly batch job that tracks processed files across runs, not a per-file event-driven trigger.

How to eliminate wrong answers

Option A is wrong because S3 Event Notifications trigger the job per object creation, which can lead to duplicate processing if multiple files arrive within the hour and does not inherently track which files have been processed across runs. Option C is wrong because storing the last processed timestamp in DynamoDB requires custom implementation and maintenance, and it does not handle edge cases like file overwrites or partitions as reliably as Glue bookmarks. Option D is wrong because S3 Inventory provides periodic snapshots (daily or weekly) and is not designed for real-time or hourly incremental processing; filtering by last modified date in the job would require full scans and manual state management.

957
MCQeasy

A company wants to schedule a nightly batch job to copy data from an on-premises PostgreSQL database to Amazon S3. The solution must minimize operational overhead. Which AWS service should be used?

A.AWS Glue
B.AWS Data Pipeline
C.Amazon EMR
D.AWS Database Migration Service (AWS DMS) with ongoing replication
AnswerA

AWS Glue provides serverless, managed ETL with built-in scheduling and JDBC connectivity to on-premises PostgreSQL, so no infrastructure is provisioned or maintained. This meets the nightly batch copy requirement to Amazon S3 while minimising operational overhead.

Why this answer

AWS Glue is the correct choice because it provides a fully managed ETL service that can connect to on-premises PostgreSQL via JDBC, extract data, and write it to Amazon S3 with minimal configuration. Glue's built-in scheduler can run the job nightly, eliminating the need to manage servers or orchestration infrastructure, which directly meets the requirement to minimize operational overhead.

Exam trap

The trap here is that candidates often confuse AWS DMS (designed for continuous replication) with batch data movement, overlooking that DMS's ongoing replication feature is not intended for scheduled batch jobs and adds unnecessary overhead for a simple nightly copy.

How to eliminate wrong answers

Option B (AWS Data Pipeline) is wrong because, while it can copy data from on-premises databases to S3, it requires managing a task runner on-premises and has higher operational overhead compared to Glue's serverless model. Option C (Amazon EMR) is wrong because it is designed for big data processing using Hadoop/Spark clusters and introduces significant overhead for cluster management, making it overkill for a simple nightly batch copy job. Option D (AWS DMS with ongoing replication) is wrong because it is primarily intended for database migration and continuous replication, not for scheduled batch jobs; using it for nightly batch copies would incur unnecessary complexity and cost for ongoing change data capture.

958
MCQmedium

A company uses Amazon Redshift for data warehousing. The data engineering team notices that queries are slow due to high disk I/O. The team wants to improve query performance without changing the cluster configuration. Which action should the team take?

A.Increase the number of nodes in the cluster.
B.Redesign tables with appropriate sort keys and distribution styles.
C.Run the ANALYZE command to update table statistics.
D.Run the VACUUM command to reclaim disk space.
AnswerB

Sort keys reduce the blocks scanned by enabling zone-map pruning, while distribution styles co-locate joined rows on the same slice, cutting inter-node network traffic. Both target disk I/O and skew directly, satisfying the constraint of improving performance without altering cluster configuration.

Why this answer

Redesigning tables with appropriate sort keys and distribution styles directly addresses high disk I/O by minimizing data scanning and reducing data movement across nodes. Sort keys enable Redshift to skip irrelevant blocks via zone maps, while distribution styles (KEY, ALL, EVEN) optimize data locality for joins and aggregations, reducing I/O without changing cluster configuration.

Exam trap

The trap here is that candidates confuse maintenance commands (ANALYZE, VACUUM) with design changes, or think scaling out (adding nodes) is allowed when the question explicitly forbids changing cluster configuration.

How to eliminate wrong answers

Option A is wrong because increasing the number of nodes changes the cluster configuration, which the question explicitly prohibits. Option C is wrong because ANALYZE updates table statistics for the query planner but does not reduce disk I/O caused by poor data layout or data movement. Option D is wrong because VACUUM reclaims disk space from deleted rows and sorts data, but it does not fundamentally redesign tables to reduce I/O; it only maintains existing design.

959
Multi-Selecthard

A financial services company needs to share sensitive customer data with a third-party analytics firm. The data resides in an S3 bucket encrypted with an AWS KMS customer managed key. The third party has their own AWS account. Which combination of steps is required to securely share the data? (Choose TWO.)

Select 2 answers
A.Share the KMS key material with the third party
B.Update the KMS key policy to include the third-party account as a principal with kms:Decrypt permission
C.Create an IAM role in the third-party account that can be assumed by the data owner
D.Grant the third-party account access to the KMS key management
E.Configure an S3 bucket policy that grants the third-party account access to the objects
AnswersB, E

The KMS key policy must name the third-party account as a principal granted kms:Decrypt, because key policies govern who may use a customer managed key. Without this grant, cross-account decryption of the S3 objects fails regardless of bucket permissions.

Why this answer

Option B is correct because when an S3 object is encrypted with a customer managed KMS key, any principal in another AWS account must be explicitly granted kms:Decrypt on that key via the key policy (or a grant); the key policy is the primary cross-account authorization mechanism for KMS keys. Option E is correct because the third-party account also needs S3-level permission to read the objects, which is provided by a bucket policy granting actions such as s3:GetObject to the third-party principal. Together, the bucket policy authorizes the S3 data access and the KMS key policy authorizes decryption of the objects.

Option A is wrong because KMS key material is never exported or shared; the third party uses the key through KMS APIs without ever seeing the material. Option C is wrong because creating a role in the third-party account assumable by the data owner reverses the trust direction and does not grant the third party access to the data. Option D is wrong because key management (e.g., kms:CreateGrant, kms:PutKeyPolicy) is not needed and would over-privilege the third party; only kms:Decrypt is required.

960
MCQhard

A company has an AWS Glue ETL job that reads from an RDS MySQL instance and writes to S3. The security team requires that the connection to RDS be encrypted and that credentials be rotated automatically. Which configuration should be used?

A.Store the database password in an encrypted parameter in Systems Manager Parameter Store and enable SSL for the connection.
B.Use IAM database authentication for RDS and store credentials in Glue connection properties.
C.Store the password in a text file in an encrypted S3 bucket and use SSL.
D.Store the password in AWS Secrets Manager with automatic rotation enabled and configure Glue to use SSL for the connection.
AnswerD

Secrets Manager with automatic rotation directly satisfies the credential-rotation constraint, unlike static Glue connection passwords. Enabling SSL encrypts data in transit between Glue and RDS MySQL, meeting the encryption requirement. Together these address both security mandates without custom rotation logic or manual credential updates.

Why this answer

AWS Secrets Manager provides automatic rotation of RDS credentials, and AWS Glue can be configured to use SSL for an encrypted connection to RDS MySQL. Option A (Systems Manager Parameter Store) stores encrypted parameters but does not natively support automatic rotation of RDS credentials. Option B (IAM database authentication) provides authentication but does not encrypt the connection itself; SSL is still required for encryption.

Option C (encrypted S3 bucket) is not a service designed for dynamic credential management and lacks automatic rotation.

961
MCQhard

A data engineer is using Amazon Athena to query Parquet data in Amazon S3. Queries are slow and scan more data than expected. The data is partitioned by year/month/day in S3, but the AWS Glue Data Catalog table has no partition metadata. Which action will improve query performance and reduce data scanned?

A.Run MSCK REPAIR TABLE or use AWS Glue Crawler to populate partition metadata in the Data Catalog.
B.Enable Athena query result reuse and set a longer result retention period.
C.Increase the number of Athena workgroups and set a higher data usage control limit.
D.Convert the data to CSV and use a SerDe that supports compression.
AnswerA

Without partition metadata, Athena cannot prune partitions and must scan the entire table location. Running MSCK REPAIR TABLE adds partitions to the Glue Data Catalog, or a Glue Crawler can discover and register them. Once partitions are registered, Athena can skip irrelevant partitions, reducing data scanned and improving speed.

Why this answer

Athena relies on the AWS Glue Data Catalog to know which partitions exist. If the table has no partition metadata, Athena treats the entire S3 prefix as one dataset and scans all files. Registering partitions via MSCK REPAIR TABLE or a Glue Crawler enables partition pruning, so queries only read the relevant partitions.

This reduces both data scanned and query latency.

Exam trap

The trap here is assuming that file format or compression changes are needed, when the real issue is that partition metadata is absent from the Data Catalog.

962
MCQmedium

A data engineer manages an AWS Glue ETL job that reads JSON files from Amazon S3 and writes to Amazon Redshift. The job recently started failing with the error: 'Unable to find catalog table' when trying to access a table in the AWS Glue Data Catalog. The engineer confirms that the table exists in the Data Catalog and that the IAM role used by the job has glue:GetTable permissions. What is the most likely cause of this error?

A.The Glue job is running in a different AWS region than the Data Catalog.
B.The table name in the Glue job script is misspelled or uses incorrect casing.
C.The Glue job's IAM role lacks permissions to access the S3 bucket containing the table's data.
D.The Glue job's IAM role lacks permissions to access the AWS Glue Data Catalog.
AnswerB

The AWS Glue Data Catalog is case-sensitive for table names. If the job script references a table name with incorrect casing or a typo, the catalog lookup fails, producing the 'Unable to find catalog table' error. Verifying the exact table name in the Glue job script against the Data Catalog resolves this issue.

Why this answer

The error indicates that the Glue job cannot locate the specified table in the Data Catalog. Since the table exists and the IAM role has the necessary glue:GetTable permission, the most likely cause is a mismatch between the table name referenced in the job script and the actual table name in the catalog. The Data Catalog is case-sensitive, so even a minor typo or casing difference will result in this error.

Exam trap

The trap here is assuming that the error is due to missing IAM permissions when the scenario already confirms that glue:GetTable is allowed.

963
Multi-Selecthard

A data engineer is troubleshooting an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The job runs successfully but 5% of records are missing after the load. The engineer suspects data consistency issues. Which THREE actions could help diagnose and resolve the problem? (Choose THREE.)

Select 3 answers
A.Use the Redshift COPY command with a manifest file to load data.
B.Increase the number of DPUs for the Glue job.
C.Enable Glue job bookmarks to track processed files.
D.Use a staging table in Redshift with a transaction to commit.
E.Review the job's CloudWatch Logs for any error messages.
AnswersA, C, E

A manifest file explicitly lists every S3 object to load, so COPY fails loudly on missing or unreadable files rather than silently skipping them. This surfaces the 5% discrepancy and prevents partial loads, addressing the suspected consistency issue.

Why this answer

Option A is correct because using the Redshift COPY command with a manifest file explicitly lists every S3 object to be loaded, preventing silent omissions of files that can occur with prefix-based loads and making missing records easier to detect. Option C is correct because Glue job bookmarks track which S3 files have already been processed, so enabling them prevents files from being skipped or reprocessed and helps identify gaps in the input data that cause missing records. Option E is correct because reviewing the job's CloudWatch Logs surfaces errors, warnings, and skipped-record messages from Glue and Redshift, which is the primary diagnostic step for understanding why 5% of records are missing.

Option B does not belong because increasing DPUs only scales compute capacity and does not address data consistency or missing records. Option D does not belong because a staging table with a transaction ensures atomicity of the load but does not by itself diagnose or resolve records being dropped during extraction or copy.

Exam trap

The trap here is that candidates often assume performance tuning (increasing DPUs) or database-level transactions (staging tables) can fix data ingestion gaps, when the actual problem is incomplete or inconsistent file discovery from the source (S3).

964
MCQhard

A company uses Amazon S3 to store large datasets. The data engineering team needs to provide access to specific objects in the bucket to external partners using presigned URLs. Each URL should expire after 12 hours. The team wants to ensure that the presigned URLs cannot be used to access other objects in the bucket. Which approach should be taken?

A.Create an IAM role for each partner and attach a policy that grants access to specific objects.
B.Generate presigned URLs using the AWS SDK, specifying the exact object key and expiration time.
C.Use a bucket policy that allows access only from the partner's IP address range.
D.Use CloudFront signed URLs with a custom policy that restricts access to specific objects.
AnswerB

Generating presigned URLs with the AWS SDK while specifying the exact object key scopes each URL to a single object, and the 12-hour expiry parameter satisfies the required lifetime. This prevents access to other objects in the bucket.

Why this answer

Presigned URLs generated via the AWS SDK allow you to specify the exact object key and expiration time, ensuring that the URL grants access only to that specific object for the defined 12-hour period. This approach uses the secret key of the IAM user or role to sign the URL, and the signature is tied to the object key, so the URL cannot be reused to access other objects in the bucket.

Exam trap

The DEA-C01 exam often tests the distinction between presigned URLs (which are tied to a specific object key and expiration) and bucket policies or IAM roles (which grant broader access), leading candidates to overcomplicate the solution with CloudFront or IP-based restrictions when a simple SDK-generated presigned URL is sufficient.

How to eliminate wrong answers

Option A is wrong because creating an IAM role for each partner and attaching a policy that grants access to specific objects does not inherently enforce time-limited access; the role would need additional mechanisms like STS to generate temporary credentials, and it does not provide the simplicity of a single URL. Option C is wrong because a bucket policy that allows access only from the partner's IP address range would grant access to all objects in the bucket (or a broader set) rather than restricting to specific objects, and it does not provide time-limited access. Option D is wrong because CloudFront signed URLs require CloudFront distribution and custom origin setup, which adds unnecessary complexity and cost; while they can restrict access to specific objects, they are not the simplest or most direct solution for S3 presigned URLs, and the question specifically asks for presigned URLs.

965
MCQmedium

A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to an Amazon Redshift cluster. The security team requires that the data be encrypted in transit between Glue and Redshift. Which configuration should the engineer implement to meet this requirement?

A.Enable SSL in the Redshift connection by setting the sslmode parameter to 'require' in the JDBC URL used by the Glue job.
B.Use AWS Glue's built-in encryption feature by setting the --encryption-type parameter to 'ssl' in the job configuration.
C.Enable encryption on the Redshift cluster by turning on cluster encryption, which automatically encrypts data in transit.
D.Configure the Glue job to write to Redshift using the Redshift Spectrum feature, which automatically encrypts data in transit.
AnswerA

Amazon Redshift supports SSL encryption for connections. To enforce encryption in transit between Glue and Redshift, the JDBC connection must use SSL. Setting sslmode=require in the JDBC URL ensures that the connection uses SSL and fails if SSL is not available. This is the standard method to encrypt data in transit for Redshift connections from Glue.

Why this answer

To encrypt data in transit between AWS Glue and Amazon Redshift, the Glue job must establish an SSL connection to Redshift. This is done by adding sslmode=require to the JDBC URL in the Glue connection. This ensures that the connection uses SSL and rejects non-SSL connections.

Other options either address encryption at rest or are not applicable to the scenario.

Exam trap

The trap here is confusing encryption at rest (cluster encryption) with encryption in transit (SSL/TLS), and assuming that enabling cluster encryption covers in-transit encryption.

966
MCQeasy

A data engineer needs to store JSON documents that are accessed by a serverless application using AWS Lambda. The documents are frequently updated and need low latency (single-digit milliseconds) for read and write operations. Which AWS service should the engineer use?

A.Amazon DynamoDB
B.Amazon ElastiCache for Redis
C.Amazon S3 (with S3 Select)
D.Amazon RDS for MySQL
AnswerA

DynamoDB delivers consistent single-digit-millisecond latency for both reads and writes at any scale, and its serverless, fully managed design pairs directly with Lambda. The stem's frequent updates and low-latency requirement rule out S3, which offers higher and more variable latency.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value and document database that provides single-digit millisecond latency for read and write operations at any scale. It natively supports JSON documents, integrates directly with AWS Lambda via the AWS SDK, and handles frequent updates efficiently through its auto-scaling and on-demand capacity modes, making it ideal for serverless applications requiring low-latency data access.

Exam trap

The trap here is that candidates often confuse ElastiCache for Redis as a primary data store due to its low latency, overlooking that it is an in-memory cache with no built-in persistence guarantees, whereas DynamoDB provides both low latency and durable, persistent storage for JSON documents.

How to eliminate wrong answers

Option B is wrong because Amazon ElastiCache for Redis is an in-memory cache, not a durable data store; while it offers sub-millisecond latency, it is typically used for caching or session management and requires a separate persistent database to avoid data loss on node failure, making it unsuitable as the primary store for frequently updated JSON documents that must persist. Option C is wrong because Amazon S3 is an object storage service with eventual consistency for overwrite PUTS and higher latency (typically tens to hundreds of milliseconds) for read operations, and S3 Select is a server-side filtering feature that does not reduce latency for individual document reads or writes; it is not designed for frequent, low-latency updates. Option D is wrong because Amazon RDS for MySQL is a relational database that requires schema definition, does not natively store JSON as a first-class document model (though it supports JSON data type, it lacks the flexible schema and single-digit millisecond read/write performance of DynamoDB for key-value access patterns), and incurs higher operational overhead for scaling and connection management in a serverless architecture.

967
MCQmedium

A data engineer is using AWS Glue ETL to transform a large dataset in S3. The job processes 2 TB of data daily and currently runs for 6 hours. The engineer wants to reduce runtime without changing the transformation logic. What is the best approach?

A.Reduce the number of DPUs to minimize overhead.
B.Use the Spark UI to analyze bottlenecks and rewrite code.
C.Increase the number of Glue DPUs or enable auto-scaling.
D.Switch from Spark to Python shell.
AnswerC

Adding DPUs or enabling auto-scaling increases the number of concurrent executors processing partitions in parallel, so the same 2 TB workload completes faster. This reduces the six-hour runtime without altering transformation logic, satisfying the constraint of unchanged code.

Why this answer

Increasing the number of DPUs or enabling auto-scaling directly allocates more distributed processing capacity to the AWS Glue job, which reduces runtime for large datasets by parallelizing the workload across more resources. Since the transformation logic is fixed and the job is already running on Spark, adding compute capacity is the most straightforward way to speed up processing without code changes.

Exam trap

The trap here is that candidates may think reducing DPUs reduces overhead and speeds up the job, but in distributed systems, more parallelism (more DPUs) reduces runtime for large datasets, while reducing DPUs increases it.

How to eliminate wrong answers

Option A is wrong because reducing DPUs would decrease parallelism and likely increase runtime, not reduce it, as the job already takes 6 hours on 2 TB of data. Option B is wrong because using the Spark UI to analyze bottlenecks and rewriting code would change the transformation logic, which the question explicitly prohibits. Option D is wrong because switching from Spark to Python shell would remove distributed processing entirely, making the job unable to handle 2 TB of data efficiently and likely causing it to fail or run far longer.

968
MCQeasy

A data engineer needs to ingest data from an on-premises Oracle database into Amazon S3. The data volume is about 500 GB initially, with daily incremental updates of 10 GB. The pipeline must minimize operational overhead. Which AWS service should be used for the initial and incremental loads?

A.AWS Database Migration Service (DMS) with change data capture (CDC) to Amazon S3.
B.AWS Glue with a JDBC connection and incremental crawl.
C.Amazon Kinesis Data Firehose with a custom producer.
D.AWS Data Pipeline with a SQL activity and HiveCopyActivity.
AnswerA

AWS DMS with change data capture performs the initial full load and then continuously replicates ongoing changes from Oracle to Amazon S3, so incremental updates need no custom scripting or scheduled jobs, satisfying the minimal-operational-overhead requirement.

Why this answer

AWS DMS with CDC is the correct choice because it supports continuous replication from Oracle to Amazon S3 with minimal overhead. It handles both the initial 500 GB full load and ongoing 10 GB daily increments via change data capture, without requiring custom code or complex pipeline management.

Exam trap

The trap here is that candidates often choose AWS Glue for its serverless nature, but Glue's incremental crawl only updates the Data Catalog, not the data itself, and it cannot capture row-level changes from a database without full reloads.

How to eliminate wrong answers

Option B is wrong because AWS Glue with an incremental crawl is designed for cataloging schema changes, not for capturing row-level changes from a database; it would require full table scans for each incremental load, which is inefficient for 10 GB daily updates. Option C is wrong because Amazon Kinesis Data Firehose requires a custom producer to stream data from Oracle, which adds operational overhead and does not natively support CDC or initial bulk loads from a database. Option D is wrong because AWS Data Pipeline with a SQL activity and HiveCopyActivity is a legacy service that lacks native CDC support for Oracle, requiring custom scripting for incremental loads and increasing operational complexity.

969
MCQhard

A data engineer is monitoring an Amazon Kinesis Data Analytics application that processes real-time clickstream data. The application uses a Flink application with multiple operators. The engineer notices that the 'millisBehindLatest' metric is increasing steadily. Which action is MOST likely to reduce the lag?

A.Decrease the batch size in the Flink application.
B.Switch the source stream to use GZIP compression.
C.Increase the parallelism of the Flink application.
D.Increase the retention period of the Kinesis stream.
AnswerC

Raising Flink parallelism distributes the lagging operators across more subtasks, so each processes a smaller share of the incoming clickstream and drains the backlog faster. This directly addresses the steadily rising millisBehindLatest by adding processing capacity, provided the source shards and downstream sinks can absorb the extra throughput.

Why this answer

Increasing the parallelism of the Flink application allows more parallel subtasks to process the incoming data, thereby increasing throughput and reducing the lag (millisBehindLatest). This is the most direct way to scale the application to handle higher load. Other options do not address the root cause of increasing lag.

Exam trap

DEA-C01 often tests the interpretation of millisBehindLatest and the appropriate scaling action, and candidates may mistakenly choose options that alter data format or retention rather than addressing processing capacity.

How to eliminate wrong answers

Option A is wrong because decreasing batch size may reduce latency but does not increase processing capacity; it could even reduce throughput. Option B is wrong because GZIP compression on the source stream reduces network bandwidth but adds CPU overhead for decompression, potentially worsening lag. Option D is wrong because increasing retention period only affects how long data is stored, not how fast it is processed.

970
Multi-Selectmedium

A data engineer is using Amazon EMR to process large datasets. The cluster uses a mix of Spot Instances and On-Demand Instances. The engineer wants to reduce costs while ensuring the job can complete even if Spot Instances are reclaimed. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Enable Instance Fleets to use multiple instance types for Spot Instances.
B.Use only On-Demand Instances for all nodes.
C.Use Spot Instances for core nodes to reduce cost.
D.Enable termination protection for the cluster.
E.Use a task instance group with Spot Instances for non-critical processing tasks.
AnswersA, E

Instance Fleets let EMR select from many Spot instance types and availability zones, so a reclamation on one pool is absorbed by others. This reduces Spot interruption risk while the job continues, complementing On-Demand capacity.

Why this answer

Option A is correct because EMR Instance Fleets let you specify multiple instance types and subnets for Spot capacity, so if one Spot pool is reclaimed or unavailable, EMR can provision from another pool, improving the chance the job finishes despite Spot interruptions. Option E is correct because task nodes perform only data-processing work and hold no HDFS data, so placing them in a Spot-based task instance group lets you capture Spot savings while core nodes (which store HDFS data) remain protected and the job can still complete if task Spot capacity is reclaimed. Option B is wrong because using only On-Demand Instances eliminates Spot savings and does not address cost reduction.

Option C is wrong because core nodes hold HDFS data; using Spot for core nodes risks data loss and job failure when Spot capacity is reclaimed. Option D is wrong because termination protection only prevents accidental cluster termination and does not help the job survive Spot reclamation or reduce costs.

Exam trap

DEA-C01 often tests the core-vs-task node distinction, and candidates who put Spot on core nodes to save money ignore that core nodes hold HDFS data, making the job vulnerable to reclamation.

971
MCQhard

A data pipeline uses AWS DMS to replicate data from an on-premises Oracle database to Amazon S3 in Parquet format. The pipeline has been running successfully for months, but recently the DMS task status shows 'failed' with the error: 'The source database is running out of archive log space.' Which action should the engineer take to prevent this error?

A.Configure multiple target S3 buckets to distribute the load.
B.Increase the amount of archive log space or reduce the log retention period on the source Oracle database.
C.Enable automatic log archiving on the DMS replication instance.
D.Increase the memory allocation for the DMS replication instance.
AnswerB

Increasing archive log space or shortening retention directly addresses the Oracle constraint: redo logs cannot be archived once the destination fills, so the DMS task fails. DMS reads redo/archive logs via LogMiner to capture ongoing changes, so the source must retain sufficient archive log capacity for replication to continue.

Why this answer

The error 'source database is running out of archive log space' originates on the Oracle source, not on DMS. Oracle must retain archive logs until DMS (via LogMiner or Binary Reader) has consumed them; if DMS lags or the retention is too long, archive log space fills up. The fix is to increase archive log space or reduce the retention period so Oracle can purge logs after DMS reads them.

Exam trap

DEA-C01 often tests whether candidates correctly attribute source-side errors to the source database rather than the DMS replication instance — the trap is assuming DMS configuration changes (memory, buckets) will fix an Oracle archive log issue.

How to eliminate wrong answers

Option A is wrong because adding target S3 buckets does not affect source-side archive log consumption — the bottleneck is on the Oracle source. Option C is wrong because DMS does not perform log archiving on the source; Oracle manages its own archive logs, and DMS only reads them. Option D is wrong because replication instance memory affects DMS processing capacity, not the source database's archive log space.

972
MCQmedium

A data engineer is troubleshooting a nightly ETL job that reads data from an RDS MySQL instance and writes to an S3 bucket in Parquet format. The job runs on an EMR cluster and uses PySpark. Recently, the job started failing with 'OutOfMemoryError' in the executor logs. The data volume has grown 30% in the last month. Which is the MOST efficient solution to resolve this issue without changing the code?

A.Change the RDS instance to a larger size to reduce load.
B.Switch the ETL job to use AWS Glue with a larger WorkerType.
C.Increase the executor memory and memoryOverhead in the Spark configuration.
D.Increase the number of core nodes in the EMR cluster.
AnswerC

Raising executor memory and memoryOverhead directly addresses the heap exhaustion causing the OutOfMemoryError, since the 30% data growth increased per-partition working set. This satisfies the no-code-change constraint because both are Spark configuration properties applied at submit time, requiring no PySpark edits.

Why this answer

The OutOfMemoryError in executors indicates insufficient memory per executor to handle the increased data volume. Increasing 'spark.executor.memory' and 'spark.executor.memoryOverhead' directly addresses this by providing more heap and off-heap memory without any code changes. Option A is wrong because the RDS instance size does not affect executor memory; the bottleneck is in Spark processing.

Option B is wrong because switching to AWS Glue would require code changes and may still need memory tuning, making it less efficient. Option D is wrong because adding core nodes increases parallelism but does not increase memory per executor, so the OOM could still occur.

973
Multi-Selecthard

Which THREE factors should be considered when choosing between Amazon Kinesis Data Streams and Amazon Kinesis Data Firehose for real-time data ingestion? (Choose three.)

Select 3 answers
A.The need for a fully managed delivery destination
B.Whether the application requires custom data processing logic
C.The ability to compress data before storage
D.The latency requirements for data delivery
E.The maximum throughput supported per shard
AnswersA, B, D

Firehose can directly deliver to S3, Redshift, etc., while Streams requires a consumer.

Why this answer

Amazon Kinesis Data Firehose is a fully managed service that automatically delivers streaming data to destinations like Amazon S3, Amazon Redshift, Amazon OpenSearch Service, and Splunk, making it ideal when you need a fully managed delivery destination without managing the ingestion pipeline. In contrast, Amazon Kinesis Data Streams requires you to build and manage consumers to process and deliver data, so if you need a fully managed destination, Firehose is the correct choice.

Exam trap

The trap here is that candidates often assume compression or throughput limits are unique to one service, but both services support compression and throughput is a scaling detail of Streams, not a direct comparison factor for choosing between the two.

974
MCQeasy

A company wants to ingest real-time streaming data from thousands of IoT devices into AWS for immediate processing. Which service is designed for ingesting large volumes of streaming data with low latency?

A.Amazon Simple Storage Service (S3)
B.AWS Database Migration Service (DMS)
C.Amazon Simple Queue Service (SQS)
D.Amazon Kinesis Data Streams
AnswerD

Amazon Kinesis Data Streams ingests thousands of concurrent device feeds with sub-second latency, satisfying the immediate-processing constraint. Unlike Kinesis Data Firehose, which micro-batches to destinations, Data Streams delivers records to consumers within milliseconds, so downstream analytics act on events as they arrive rather than after buffering.

Why this answer

Amazon Kinesis Data Streams is purpose-built for real-time streaming data ingestion at scale. Option A (S3) is object storage, not designed for streaming ingestion. Option B (DMS) is for database migration, not real-time data streaming.

Option C (SQS) is a message queue service that is pull-based and not optimized for high-throughput streaming with low latency. Only option D (Kinesis Data Streams) meets all requirements for ingesting large volumes of streaming data with low latency.

975
MCQhard

A company runs a critical PostgreSQL database on Amazon RDS. The database experiences high read latency during peak hours. The data engineer needs to reduce read latency with minimal changes to the application. Which solution is MOST effective?

A.Delete unused indexes to improve query performance.
B.Enable Multi-AZ deployment for automatic failover.
C.Increase the DB instance class to a larger size with more memory.
D.Create a read replica of the RDS instance and redirect read queries to it.
AnswerD

Read replicas offload read traffic to a separate RDS instance via PostgreSQL's asynchronous streaming replication, directly reducing contention on the primary and cutting read latency. Redirecting read queries requires only connection-string changes, satisfying the stem's minimal-application-change constraint. Unlike Multi-AZ standby, replicas serve reads.

Why this answer

A read replica is a separate RDS instance that receives asynchronous replication from the primary. By redirecting read-only queries to the replica, you offload read traffic from the primary, directly reducing read latency during peak hours. This requires minimal application changes—typically just updating the connection string for read operations.

Multi-AZ and larger instance classes do not provide a dedicated read-scaling endpoint, and deleting indexes would harm performance.

Exam trap

DEA-C01 often tests the misconception that Multi-AZ or scaling up the primary instance improves read performance, when in fact only read replicas provide read scaling with minimal application changes.

How to eliminate wrong answers

Option A is wrong because deleting unused indexes may reduce write overhead but does not address read latency; in fact, removing indexes that are occasionally used can force full table scans, increasing read latency. Option B is wrong because Multi-AZ provides high availability through a standby that does not serve read traffic; it does not reduce read latency on the primary. Option C is wrong because scaling up the primary instance may temporarily improve performance, but it does not isolate read workloads and can be more expensive and disruptive than adding a read replica.

Page 12

Page 13 of 18

Page 14