Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 301–375

1321 questions total · 18pages · All types, answers revealed

Page 4

Page 5 of 18

Page 6
301
MCQeasy

A company's analytics team needs a petabyte-scale, fully managed data warehouse that supports standard SQL, columnar storage, and massively parallel query execution, and it must integrate with existing business intelligence tools with minimal operational effort. Which AWS service should the data engineer choose?

A.Amazon Redshift
B.AWS Database Migration Service
C.Amazon ElastiCache for Redis
D.Amazon DynamoDB
AnswerA

Amazon Redshift is a fully managed, petabyte-scale data warehouse that uses columnar storage on a cluster of nodes with massively parallel processing, and it speaks standard SQL that BI tools connect to through JDBC and ODBC drivers. It matches every stated requirement, including minimal operational effort, because AWS handles provisioning, patching, and backups of the cluster.

Why this answer

The workload calls for a managed analytical database with columnar storage, massively parallel execution, and standard SQL connectivity for BI tools. Amazon Redshift delivers all three as a managed service, so the team avoids building and operating its own warehouse. The other services target transactional NoSQL access, in-memory caching, or data movement rather than petabyte-scale SQL analytics.

Exam trap

The trap here is being drawn to any highly scalable AWS data store and overlooking that only a columnar, MPP SQL warehouse satisfies the analytics and BI requirements.

302
MCQmedium

A company uses Amazon EMR to run Spark jobs on data stored in S3. After upgrading the EMR cluster to a new release, one of the Spark jobs fails with 'OutOfMemoryError' in the executor. Which configuration change is MOST likely to resolve this issue?

A.Increase the number of core nodes in the EMR cluster.
B.Decrease spark.sql.shuffle.partitions to reduce overhead.
C.Increase spark.driver.memory in the Spark configuration.
D.Increase spark.executor.memory to allocate more memory per executor.
AnswerD

Spark executor memory is set by spark.executor.memory; the OutOfMemoryError arises because the new EMR release changed default executor sizing, leaving too little heap. Raising this value directly expands executor heap, satisfying the job's memory constraint without altering cluster hardware.

Why this answer

Increasing spark.executor.memory directly allocates more memory per executor, addressing the OutOfMemoryError in the executor. Option A (increasing the number of core nodes) adds more cluster capacity but does not increase the memory available to individual executors, so it may not resolve the OOM if the existing executors are already memory-constrained. Option B (decreasing spark.sql.shuffle.partitions) reduces the number of shuffle partitions, which can increase the size of each partition and potentially cause more memory pressure, not less.

Option C (increasing spark.driver.memory) only helps the driver process, not the executor, so it does not fix executor OOM errors.

303
MCQmedium

A data engineer runs an AWS Glue ETL job that reads CSV files from an Amazon S3 bucket, applies transformations, and writes Parquet output to another S3 bucket. The job fails with the error 'AnalysisException: Unable to infer schema for CSV. It must be specified manually.' The CSV files are stored with a header row, and the job's script uses the default Glue DynamicFrame reader without specifying format options. What is the MOST likely cause of the failure?

A.The CSV files are compressed with gzip, and Glue cannot infer schema from compressed CSV files.
B.The CSV files have inconsistent or malformed data that prevents Glue from inferring a schema, or the header option is not set correctly.
C.The S3 bucket containing the CSV files does not have the correct bucket policy allowing AWS Glue to read objects.
D.The Glue job's IAM role lacks permissions to read the CSV files from the S3 bucket.
AnswerB

Glue's schema inference can fail if CSV files have inconsistent columns, missing headers, or if the 'withHeader' option is not set to true. In this scenario, the header row exists but the default reader may not treat it as a header, leading to type conflicts. Setting 'withHeader' to true and ensuring consistent data resolves the error.

Why this answer

The error 'Unable to infer schema for CSV' occurs when AWS Glue cannot determine column names and types from the source data. This often happens when the header option is not enabled or when the CSV data has inconsistencies such as varying column counts or mixed data types. Ensuring the reader is configured with 'withHeader' set to true and that data is well-formed allows Glue to infer the schema correctly.

Exam trap

The trap here is assuming that a schema inference error is caused by permissions or compression, when it actually stems from data formatting or missing header configuration.

304
MCQmedium

A data engineer needs to migrate an on-premises Apache Hadoop cluster to AWS. The cluster stores data in HDFS and runs MapReduce jobs. The company wants to minimize operational overhead and leverage serverless technologies where possible. Which AWS service should the data engineer use to replace HDFS storage?

A.Amazon EBS
B.Amazon EMR
C.Amazon S3
D.Amazon Redshift
AnswerC

Amazon S3 provides durable, serverless object storage that replaces HDFS without cluster management, satisfying the minimise-operational-overhead constraint. MapReduce workloads migrate to Amazon EMR or Athena, while S3 becomes the decoupled storage layer, eliminating NameNode and DataNode administration entirely.

Why this answer

Amazon S3 is the correct replacement for HDFS because it provides highly durable, scalable, and serverless object storage that can be used as the primary storage layer for Amazon EMR. Unlike HDFS, S3 decouples storage from compute, eliminating the need to manage cluster storage and allowing jobs to run on ephemeral clusters, which minimizes operational overhead. S3 integrates with EMR via the EMR File System (EMRFS), enabling MapReduce jobs to read/write data directly from S3 as if it were HDFS.

Exam trap

The trap here is that candidates confuse Amazon EMR (a compute service) with a storage service, assuming it replaces HDFS, when in fact EMR can use either HDFS or S3 for storage, and the question explicitly asks for the storage replacement.

How to eliminate wrong answers

Option A is wrong because Amazon EBS provides block-level storage volumes attached to EC2 instances, which is not serverless and requires manual management of volume size, snapshots, and replication; it also ties storage to a specific compute instance, defeating the purpose of decoupling storage from compute for a Hadoop migration. Option B is wrong because Amazon EMR is a managed big data platform that runs MapReduce jobs, not a storage service; it can use HDFS or S3 for storage, but the question specifically asks for a replacement of HDFS storage, not the compute framework. Option D is wrong because Amazon Redshift is a fully managed data warehouse optimized for SQL-based analytics and structured data, not a general-purpose distributed file system for Hadoop workloads; it does not support HDFS semantics or MapReduce jobs natively.

305
MCQeasy

A data engineer has set up an AWS Lambda function that processes files uploaded to an S3 bucket. The function is triggered by S3 event notifications. However, the function is not being invoked when a file is uploaded. The engineer checks the Lambda function's CloudWatch Logs and finds no execution logs. What should the engineer check FIRST?

A.Check the Lambda function's code for errors.
B.Verify that the Lambda function's IAM role has permissions to read from S3.
C.Verify that the S3 bucket has an event notification configured for the Lambda function.
D.Check if the Lambda function is attached to a VPC.
AnswerC

Without an S3 event notification targeting the Lambda function, uploads never generate an invocation, so no CloudWatch execution logs appear. Confirming the notification configuration first isolates whether the trigger exists before investigating permissions or code.

Why this answer

The symptom is that the Lambda function is not invoked at all and there are no CloudWatch execution logs — meaning the function never ran. The first thing to verify is whether the S3 bucket actually has an event notification configured to trigger the Lambda function, since without that configuration no invocation occurs. This is the most direct cause of 'no invocation, no logs.'

Exam trap

DEA-C01 often tests the misconception that a Lambda failure is always a code or IAM issue — candidates must remember that 'no invocation, no logs' points to the trigger configuration (S3 event notification) being missing or misconfigured.

How to eliminate wrong answers

Option A is wrong because code errors would only manifest after the function is invoked — and the absence of execution logs proves the function never ran. Option B is wrong because IAM permissions on the Lambda role matter only after invocation; they cannot prevent the S3 event from triggering the function. Option D is wrong because VPC attachment affects network access from the function, not whether the function is invoked — and again, no logs means no invocation.

306
MCQmedium

A data engineer needs to transform CSV files arriving in an S3 bucket into Parquet format and store them in another S3 bucket. The transformation is simple and on-demand, triggered by data arrival. Which solution is the MOST cost-effective and requires the least operational overhead?

A.Use Amazon EMR with Spark streaming
B.Use Amazon Athena to create a new table with Parquet format
C.Use AWS Glue ETL jobs scheduled to run every hour
D.Use S3 Events to trigger an AWS Lambda function that transforms the data
AnswerD

S3 event notifications invoke Lambda directly on object arrival, so no cluster runs between executions and no polling infrastructure is needed. Lambda's per-invocation billing suits sporadic, on-demand CSV-to-Parquet conversion, and the service manages scaling and patching, satisfying the minimal operational overhead constraint.

Why this answer

Using S3 Events to trigger an AWS Lambda function is the most cost-effective and operationally lightweight solution for simple, on-demand CSV-to-Parquet transformations. Lambda scales automatically with each S3 PUT event, incurs no idle cost, and requires no cluster management, making it ideal for event-driven, low-volume transformations.

Exam trap

The DEA-C01 exam often tests the misconception that AWS Glue is always the best choice for ETL, but for simple, event-driven transformations with minimal overhead, Lambda is more cost-effective and operationally simpler than Glue's managed Spark environment.

How to eliminate wrong answers

Option A is wrong because Amazon EMR with Spark streaming introduces significant operational overhead (cluster provisioning, scaling, and management) and cost (even with auto-scaling, you pay for running instances) for a simple, on-demand transformation that does not require real-time streaming. Option B is wrong because Amazon Athena cannot transform data into Parquet format; it is a query engine that can read from and write to Parquet tables via CTAS statements, but it does not provide a direct, event-driven transformation trigger and incurs per-query costs that can be higher than Lambda for frequent small files. Option C is wrong because AWS Glue ETL jobs scheduled every hour introduce unnecessary latency (up to 1 hour delay) and cost (minimum billing per DPU hour) for an on-demand workload triggered by data arrival, and the scheduled polling approach is less efficient than event-driven invocation.

307
Drag & Dropmedium

Order the steps to troubleshoot a failed AWS Glue job that reads from JDBC and writes to S3.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Start with logs to identify errors, then check connectivity, IAM permissions, test connection, and review script.

308
MCQeasy

A data engineer notices that a nightly AWS Glue ETL job has been failing for the past three days with the error 'Unable to locate credentials'. The job uses an IAM role for execution. What is the most likely cause of this error?

A.The IAM role does not have an access key attached.
B.The S3 bucket name in the job parameters is misspelled.
C.The IAM role's trust policy does not include glue.amazonaws.com as a trusted entity.
D.The JDBC connection string contains an incorrect password.
AnswerC

Glue assumes its execution role via AWS Security Token Service, which validates the role's trust policy. If glue.amazonaws.com is absent as a trusted principal, AssumeRole is denied, producing the 'Unable to locate credentials' error rather than a permissions failure.

Why this answer

The error 'Unable to locate credentials' indicates that the AWS Glue job cannot obtain AWS credentials to authenticate API calls. Since the job uses an IAM role for execution, the most likely cause is that the trust policy of that IAM role does not include 'glue.amazonaws.com' as a trusted entity. Without this trust relationship, AWS Glue cannot assume the role and thus has no credentials to sign requests.

Exam trap

AWS often tests the distinction between IAM role trust policies (who can assume the role) and IAM role permission policies (what actions the role can perform), and candidates mistakenly focus on permission policies when the error is about credential acquisition.

How to eliminate wrong answers

Option A is wrong because IAM roles do not use access keys; they use temporary security credentials obtained via the AWS Security Token Service (STS). Option B is wrong because a misspelled S3 bucket name would cause a 'NoSuchBucket' or 'Access Denied' error, not a credentials-related error. Option D is wrong because an incorrect JDBC password would result in a connection failure or authentication error from the database, not an 'Unable to locate credentials' error from AWS.

309
Multi-Selecthard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer notices that some queries are waiting in the queue for a long time, and the WLM (Workload Management) configuration is set to auto. The engineer wants to implement manual WLM to improve query throughput and ensure that short-running queries are not blocked by long-running ones. Which TWO actions should the engineer take to achieve this? (Choose two.)

Select 2 answers
A.Set the WLM timeout for the long-running query queue to automatically terminate queries that exceed a specified time.
B.Create separate WLM queues for short-running and long-running queries, and assign appropriate memory percentages to each queue.
C.Use query monitoring rules (QMR) to log queries that exceed resource thresholds and send alerts.
D.Increase the number of nodes in the Redshift cluster to provide more memory and CPU resources.
E.Enable concurrency scaling to automatically add cluster capacity for bursts of queries.
AnswersA, B

Setting a WLM timeout on the long-running queue prevents those queries from monopolizing resources indefinitely. When a query exceeds the timeout, it is terminated, freeing up resources for other queries. This helps maintain throughput and ensures that short queries are not starved. It is a key configuration in manual WLM to manage runaway queries.

Why this answer

Manual WLM with separate queues for short and long queries, along with WLM timeouts for long queries, directly addresses the issue of short queries being blocked. Separate queues provide resource isolation, while timeouts prevent long queries from consuming resources indefinitely. These two actions together improve throughput and ensure responsiveness during peak hours.

Exam trap

The trap here is confusing concurrency scaling or QMR as solutions for queue wait times, when they do not provide the necessary workload isolation and resource management.

310
Multi-Selectmedium

A data engineering team uses AWS Glue to extract, transform, and load (ETL) data from Amazon RDS for MySQL to Amazon S3. The job runs daily and processes incremental data. The team notices that the job is taking longer than expected. Which TWO actions can improve the job performance? (Choose two.)

Select 2 answers
A.Change the worker type to Standard (single node).
B.Use pushdown predicates to filter data at the source.
C.Add more transformations to the ETL script to clean data.
D.Increase the number of DPUs for the Glue job.
E.Disable compression on the output data to reduce CPU usage.
AnswersB, D

Pushdown predicates translate filter conditions into SQL WHERE clauses executed by RDS for MySQL, so only matching incremental rows are read over JDBC. Less data is transferred and processed in Glue, directly reducing the job's runtime.

Why this answer

Option B is correct because pushdown predicates let AWS Glue push filtering logic down to the source RDS for MySQL database, so only the required incremental rows are read over JDBC instead of the entire table, reducing I/O and shuffle work in the job. Option D is correct because increasing the number of DPUs adds more Apache Spark executors and parallel task slots, which improves throughput for a large, daily incremental ETL workload that is currently resource-bound. Option A is not appropriate because switching to a Standard single-node worker removes distributed processing and would generally slow the job rather than improve performance.

Option C is not appropriate because adding more transformations increases CPU and memory work in the ETL script, which would make the job slower, not faster. Option E is not appropriate because disabling output compression increases the volume of data written to Amazon S3 and read downstream, raising I/O and cost rather than improving job performance.

Exam trap

DEA-C01 often tests the misconception that more transformations or disabling compression improve ETL speed, when in fact pushdown predicates and additional DPUs are the canonical performance levers.

311
MCQhard

A data engineer is configuring an AWS Lake Formation permissions model for a data lake in Amazon S3. Analysts must query a table through Amazon Athena and see only rows where the 'region' column equals 'EU'. The engineer has already registered the S3 location with Lake Formation and created the table in the AWS Glue Data Catalog. Which action should the engineer take to enforce the row-level restriction?

A.Create an IAM policy that allows Athena StartQueryExecution only when the query string contains region='EU'.
B.Create an S3 bucket policy that denies GetObject for objects whose keys do not start with 'EU/'.
C.Create a Lake Formation data filter that includes the expression region='EU' and grant SELECT on the table with that data filter to the analysts' IAM role.
D.Create a separate Athena workgroup for the analysts and configure the workgroup to append a WHERE clause to every query.
AnswerC

Lake Formation data filters allow row-level and cell-level security by attaching a filter expression to a table resource. When you grant SELECT with a data filter, Athena queries automatically include the filter condition, so analysts only see rows where region equals 'EU'. This is the native Lake Formation mechanism for row-level access control without modifying the underlying data.

Why this answer

Lake Formation data filters are the correct mechanism for row-level security. By creating a data filter with the expression region='EU' and granting SELECT with that filter, the analyst's queries through Athena automatically receive the filter condition. This enforcement happens at the Lake Formation permission layer, so it applies consistently regardless of how the query is written, and the underlying data remains unchanged.

Exam trap

The trap here is assuming that IAM policies or S3 bucket policies can enforce row-level filtering, when only Lake Formation data filters operate at that granularity.

312
MCQmedium

A data engineer stores application logs in an Amazon S3 bucket. Compliance requires that log objects be retained for seven years and that they cannot be deleted or overwritten by any user, including the account root user, during that period. The engineer must configure the bucket to enforce this. Which combination of settings should the engineer apply?

A.Enable S3 Object Lock in governance mode and grant the s3:BypassGovernanceRetention permission only to the root user.
B.Enable S3 Versioning and attach a bucket policy that denies s3:DeleteObject to all principals except the root user.
C.Enable S3 Object Lock in compliance mode with a default retention period of seven years on the bucket.
D.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Deep Archive after one day and expire them after seven years.
AnswerC

S3 Object Lock in compliance mode prevents any user, including the root user, from overwriting or deleting a protected object version until the retention period expires. Setting a default retention of seven years on the bucket applies the rule automatically to new objects, satisfying the immutability and retention requirements.

Why this answer

S3 Object Lock provides write-once-read-many protection at the object version level. Compliance mode is the strictest setting: no principal, including the account root user, can delete or overwrite a locked version before retention expires. Applying a bucket default retention of seven years automatically protects newly uploaded log objects without per-object calls, meeting the compliance requirement.

Exam trap

The trap here is believing that a bucket policy or governance mode can stop the root user, when only compliance-mode Object Lock removes all override capability.

313
MCQhard

A company has a requirement to store audit logs for 7 years for compliance. The logs are stored in S3 and must be immutable. Which S3 feature should be used?

A.Use a bucket policy that denies s3:DeleteObject
B.Enable MFA Delete on the bucket
C.Enable S3 Versioning and set a lifecycle policy
D.Enable S3 Object Lock in compliance mode
AnswerD

S3 Object Lock in compliance mode enforces a write-once-read-many (WORM) retention period that no user, including the root account, can shorten or bypass. This directly satisfies the seven-year immutability requirement, unlike governance mode, which privileged users can override.

Why this answer

S3 Object Lock in compliance mode prevents objects from being deleted or overwritten for a specified retention period, ensuring immutability for compliance. Option A is wrong because a bucket policy denying s3:DeleteObject does not prevent overwrites or other actions that could modify the object. Option B is wrong because MFA Delete adds a protection layer but can be bypassed by the root user and does not prevent overwrites.

Option C is wrong because versioning and lifecycle policies do not inherently prevent deletion or overwrite of all versions; they only manage object versions and transitions. Option D is correct.

314
MCQhard

A company uses Amazon S3 to store sensitive financial data. The security team requires that all objects be encrypted at rest using AWS KMS with a customer-managed key. Additionally, they want to audit all KMS decrypt calls for compliance. Which configuration should be used to meet these requirements?

A.Enable default encryption on the bucket with SSE-KMS using an AWS managed key.
B.Use SSE-S3 with a bucket policy that denies uploads without encryption.
C.Use SSE-KMS with a customer-managed KMS key and enable CloudTrail data events for the key.
D.Use SSE-C with client-managed keys and log S3 API calls.
AnswerC

SSE-KMS with a customer-managed key satisfies the encryption-at-rest requirement, while CloudTrail data events record KMS Decrypt API calls against that key, delivering the required audit trail. Default CloudTrail management events alone would not capture per-object decrypt activity.

Why this answer

SSE-KMS with a customer-managed key allows the company to control the encryption key lifecycle and meet the requirement for customer-managed keys. Enabling CloudTrail data events for the KMS key captures all decrypt API calls, providing the necessary audit trail for compliance.

Exam trap

The trap here is that candidates may confuse enabling default encryption on the bucket (which can use SSE-KMS) with the need for a customer-managed key and CloudTrail data events, or they may think SSE-S3 or SSE-C can satisfy the audit requirement without KMS-specific logging.

How to eliminate wrong answers

Option A is wrong because it uses an AWS managed key, not a customer-managed key, so the security team cannot control key rotation or access policies. Option B is wrong because SSE-S3 uses server-side encryption with S3-managed keys, which does not provide customer-managed key control, and the bucket policy only enforces encryption, not auditing of decrypt calls. Option D is wrong because SSE-C requires the client to manage the encryption keys, which does not meet the requirement for AWS KMS, and logging S3 API calls alone does not capture KMS decrypt events.

315
MCQeasy

A company stores sensitive data in Amazon S3 and requires that all data in transit between on-premises applications and S3 be encrypted. The applications use the AWS SDK to upload and download objects. Which configuration should the data engineer implement to enforce encryption in transit?

A.Configure the S3 bucket policy to deny requests that do not use the aws:SecureTransport condition key.
B.Enable default encryption on the S3 bucket using SSE-S3.
C.Enable S3 Transfer Acceleration on the bucket to ensure data is encrypted during transfer.
D.Use AWS Certificate Manager (ACM) to provision a TLS certificate for the S3 bucket.
AnswerA

The aws:SecureTransport condition key checks whether the request was sent over HTTPS. By adding a bucket policy that denies requests where aws:SecureTransport is false, you enforce that all data in transit uses TLS encryption. This is a standard and effective method to enforce encryption in transit for S3.

Why this answer

To enforce encryption in transit for S3, the bucket policy should deny any request that does not use HTTPS, which is indicated by the aws:SecureTransport condition key being false. This ensures that all data transfers between on-premises applications and S3 are encrypted using TLS. Other options address encryption at rest or performance, not in-transit encryption.

Exam trap

The trap here is confusing encryption at rest (SSE-S3) with encryption in transit, which requires enforcing HTTPS via bucket policies.

316
MCQeasy

A startup is building a mobile application that requires a database to store user profiles and preferences. The database must scale automatically with minimal administration. Which AWS service should they use?

A.Amazon Redshift
B.Amazon Aurora
C.Amazon DynamoDB
D.Amazon RDS for PostgreSQL
AnswerC

DynamoDB is fully managed and serverless, scaling throughput automatically via on-demand capacity with no server provisioning or patching. This directly satisfies the stem's constraints of automatic scaling and minimal administration for storing user profiles and preferences.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value and document database that delivers single-digit millisecond performance at any scale. It supports automatic scaling of throughput capacity and storage with no downtime, making it ideal for a mobile application that requires minimal administrative overhead. The serverless, pay-per-request billing model aligns perfectly with the startup's need for automatic scaling and low operational burden.

Exam trap

The trap here is that candidates often choose a relational database like Aurora or RDS because they assume user profiles require complex joins or ACID transactions, but DynamoDB's single-table design and conditional updates can handle most mobile app patterns with simpler, more scalable operations.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a petabyte-scale data warehouse optimized for complex analytical queries, not for transactional user profile storage, and it requires manual scaling and cluster management. Option B is wrong because Amazon Aurora is a relational database that, while offering some auto-scaling for storage, still requires manual provisioning of compute resources and is not as fully serverless as DynamoDB for this use case. Option D is wrong because Amazon RDS for PostgreSQL is a managed relational database but requires manual scaling of instance size and storage, and does not offer the same level of automatic, seamless scaling as DynamoDB for a mobile app's unpredictable workload.

317
Multi-Selecthard

Which TWO of the following are best practices for Amazon Redshift table design? (Choose TWO.)

Select 2 answers
A.Choose sort keys based on query patterns
B.Use INSERT statements for large data loads
C.Avoid compression encoding to reduce CPU overhead
D.Specify distribution keys to minimize data movement
E.Set distribution style to ALL for all tables
AnswersA, D

Sort keys determine the order in which rows are stored on disk, so choosing them from actual query filter and join patterns lets Redshift skip irrelevant blocks via zone maps, reducing I/O. This directly satisfies the best-practise requirement for table design.

Why this answer

Option A is correct because choosing sort keys based on query patterns allows Amazon Redshift to use zone maps and skip scanning irrelevant blocks, dramatically reducing I/O for range-filtered and ordered queries. Option D is correct because specifying an appropriate distribution key colocates matching rows on the same compute node slice, so joins and aggregations on that key avoid costly data redistribution across the network. Option B is wrong because large data loads should use COPY (ideally from Amazon S3) rather than INSERT, which is slow and row-by-row.

Option C is wrong because compression encoding is a best practice that reduces storage and I/O; Redshift uses columnar compression with minimal CPU penalty, so avoiding it is counterproductive. Option E is wrong because distribution style ALL replicates the entire table to every node, which is only suitable for small dimension tables and causes excessive storage and load overhead on large tables.

Exam trap

DEA-C01 often tests whether candidates confuse distribution styles (ALL vs. KEY vs. EVEN) and mistakenly believe that ALL is a safe default, when it actually causes severe storage and performance penalties for large tables.

318
MCQeasy

A data engineer needs to schedule a daily AWS Glue job that extracts data from Amazon S3 and loads it into Amazon Redshift. The engineer wants to ensure the job runs at 2:00 AM UTC every day and can be monitored for failures. What is the simplest way to achieve this?

A.Use AWS Glue triggers to schedule the job with a cron expression.
B.Create an AWS Lambda function that triggers the Glue job using a cron expression in Amazon EventBridge.
C.Use Amazon EventBridge Scheduler to invoke the Glue job directly.
D.Create an AWS Step Functions state machine with a Wait state and a Glue job task.
AnswerA

Correct. AWS Glue triggers allow you to schedule jobs using cron expressions. You can create a trigger of type SCHEDULED, set the cron expression to '0 2 * * ? *' for 2:00 AM UTC daily, and attach it to the Glue job. This is the simplest and most native way to schedule Glue jobs, with built-in monitoring through Glue job run metrics and CloudWatch.

Why this answer

AWS Glue triggers are the native scheduling mechanism for Glue jobs. Creating a scheduled trigger with a cron expression allows the job to run automatically at specified times. This approach is simple, requires no additional services, and provides monitoring through Glue job run statuses and CloudWatch metrics.

The cron expression '0 2 * * ? *' schedules the job for 2:00 AM UTC daily.

Exam trap

The trap here is overcomplicating the solution by involving additional services like Lambda or Step Functions when AWS Glue provides a built-in scheduling feature.

319
MCQmedium

A data engineer is designing a pipeline that ingests JSON logs from an application into Amazon S3. The logs contain a timestamp field. The pipeline must partition the data by date in S3 (e.g., year=2024/month=10/day=01). Which approach minimizes transformation effort?

A.Use Amazon Kinesis Data Firehose with dynamic partitioning
B.Use AWS Glue crawlers to infer schema and create partitions
C.Use AWS Lambda to process each object and copy to the appropriate prefix
D.Use Amazon Athena to create partitions on the existing data
AnswerA

Kinesis Data Firehose dynamic partitioning extracts the timestamp field via a jq expression and writes records into year=/month=/day=/ prefixes automatically, so no downstream ETL job is needed to reorganise objects. This directly satisfies the stem's requirement to minimise transformation effort while delivering date-partitioned JSON into Amazon S3.

Why this answer

Amazon Kinesis Data Firehose with dynamic partitioning can automatically partition incoming JSON data based on the timestamp field without requiring custom transformation code. It evaluates the timestamp using a JQ expression or inline parsing, then writes records directly to S3 prefixes like year=2024/month=10/day=01. This minimizes transformation effort because the partitioning logic is configured declaratively in the Firehose delivery stream, eliminating the need for Lambda functions or post-ingestion processing.

Exam trap

The trap here is that candidates confuse metadata partitioning (e.g., using Glue crawlers or Athena) with physical partitioning in S3, assuming that catalog operations alone reorganize the data, when in fact only ingestion-time partitioning (like Firehose dynamic partitioning) creates the folder structure without extra transformation effort.

How to eliminate wrong answers

Option B is wrong because AWS Glue crawlers infer schema and create partition metadata in the Glue Data Catalog, but they do not physically reorganize data into partitioned S3 prefixes; they only add partition keys to the catalog after data is already stored. Option C is wrong because using AWS Lambda to process each object and copy it to the appropriate prefix introduces significant transformation effort, including writing custom code for parsing, partitioning logic, and handling retries, which contradicts the goal of minimizing effort. Option D is wrong because Amazon Athena can create partitions on existing data using ALTER TABLE ADD PARTITION or MSCK REPAIR TABLE, but this only updates the catalog metadata and does not physically partition the data in S3; the data remains in a flat structure, and Athena queries still scan all files unless partitions are manually created.

320
Multi-Selecthard

A company is migrating a large Oracle data warehouse to Amazon Redshift. Which THREE considerations are important for optimizing the Redshift cluster?

Select 3 answers
A.Purchasing reserved instances for the cluster.
B.Using columnar storage format.
C.Defining appropriate sort keys for the tables.
D.Applying compression encoding to columns.
E.Choosing the right distribution style (KEY, ALL, EVEN).
AnswersC, D, E

Improves query performance by reducing scans.

Why this answer

Sort keys in Amazon Redshift determine the physical order of data on disk, which directly impacts the efficiency of range-restricted queries and compression. By defining appropriate sort keys (compound or interleaved), the query optimizer can use zone maps to skip large blocks of data that don't match the filter criteria, significantly reducing the number of blocks scanned and improving query performance.

Exam trap

The trap here is that candidates confuse cost-saving measures (reserved instances) or inherent architecture features (columnar storage) with active optimization choices, when in fact only sort keys, distribution styles, and compression encoding are configurable settings that directly impact query performance in Redshift.

321
MCQmedium

A data engineer needs to allow an AWS Lambda function to access a specific AWS KMS customer managed key to decrypt data. The Lambda function assumes an IAM role. Which policy statement should be added to the KMS key policy to grant the necessary permissions with least privilege?

A.Allow the IAM role to perform kms:Decrypt on the KMS key, with a condition that the request comes from the Lambda function's VPC endpoint.
B.Allow the IAM role to perform kms:Decrypt on the KMS key, specifying the role's ARN as the principal.
C.Allow the IAM role to perform kms:* on the KMS key, specifying the role's ARN as the principal.
D.Allow the IAM role to perform kms:Decrypt on the KMS key, with a condition that the aws:PrincipalArn matches the role's ARN.
AnswerB

This statement grants the exact permission needed (kms:Decrypt) to the specific IAM role. It follows least privilege by not granting additional actions and by scoping the principal to the role. This is the correct and simplest way to allow the Lambda function to decrypt data using the KMS key.

Why this answer

The KMS key policy must grant the Lambda function's IAM role the kms:Decrypt permission. Specifying the role's ARN as the principal ensures only that role can use the key for decryption. This follows least privilege by not granting unnecessary actions.

Other options either grant excessive permissions or add unnecessary conditions that could hinder legitimate access.

Exam trap

The trap here is granting kms:* or adding unnecessary conditions instead of focusing on the specific kms:Decrypt action needed.

322
MCQmedium

A data engineer is building a data lake on Amazon S3 and needs to store JSON logs that will be queried by Amazon Athena. The engineer wants to minimize query cost and improve performance by reducing the amount of data scanned. The logs are approximately 1 KB each and arrive continuously. Which solution should the engineer implement?

A.Convert the JSON logs to Apache Parquet, partition the data by date, and store it in Amazon S3. Query with Athena.
B.Store the JSON logs in Amazon S3 Standard-Infrequent Access (S3 Standard-IA) and query them with Athena.
C.Enable S3 Transfer Acceleration on the bucket and query the JSON logs with Athena.
D.Store the JSON logs in Amazon S3 Glacier Instant Retrieval and query them with Athena.
AnswerA

Converting to columnar Parquet enables Athena to read only the columns referenced in a query, and partitioning by date allows Athena to skip irrelevant partitions. Together, these reduce the amount of data scanned, lowering query cost and improving performance, which directly meets the requirement.

Why this answer

Athena charges based on the amount of data scanned per query. Converting JSON to a columnar format like Parquet allows column pruning, so only the needed columns are read, and partitioning by date enables partition pruning, so only relevant folders are scanned. This combination minimizes data scanned, reducing cost and improving query speed.

Exam trap

The trap here is assuming that changing the S3 storage class reduces Athena query costs, when Athena cost is driven by data scanned, not by where the data is stored.

323
MCQmedium

A financial analytics company stores daily transaction records in Amazon S3 as Apache Parquet files, partitioned by year/month/day. The data engineering team queries these files with Amazon Athena. To reduce query runtime and cost, they want to apply fine-grained access control and column-level filtering without changing the files. Which solution should they use?

A.Define an AWS Lake Formation table with column-level permissions and use Lake Formation to manage access for Athena users.
B.Create an AWS Glue Data Catalog table with partition projection and use Amazon S3 Access Points for authorization.
C.Enable Amazon S3 Block Public Access and use IAM policies with condition keys to filter columns.
D.Store the data in Amazon Redshift Spectrum and use Redshift database roles to restrict column access.
AnswerA

AWS Lake Formation provides fine-grained access control at the table, column, and row level for data in the AWS Glue Data Catalog. Athena integrates with Lake Formation to enforce these permissions during queries without modifying underlying data. This directly satisfies the requirement for column-level filtering and access control.

Why this answer

AWS Lake Formation is designed to provide granular access control for data lakes built on Amazon S3 and the AWS Glue Data Catalog. It allows administrators to grant or revoke permissions at the database, table, column, and row level. Athena respects these permissions, enabling column-level filtering without data duplication or transformation.

This meets the requirement to reduce runtime and cost by limiting data access.

Exam trap

The trap here is assuming that S3 Access Points or IAM policies can enforce column-level access control for Athena queries, when in fact only Lake Formation provides that granularity.

324
MCQmedium

A data engineer is using AWS Step Functions to orchestrate an ETL workflow that includes an AWS Glue job. The Glue job occasionally fails due to transient issues, such as network timeouts. The engineer wants the Step Functions state machine to automatically retry the Glue job up to three times with exponential backoff before failing the workflow. Which Step Functions state configuration should the engineer use?

A.Configure the Glue job to have a retry policy in its job definition, such as MaxRetries: 3.
B.Use a Retry field with MaxAttempts: 3 and BackoffRate: 2.0 in the task state that invokes the Glue job.
C.Use a Catch field with ErrorEquals: ["States.ALL"] and Next: "RetryState" to redirect to a retry state.
D.Set the TimeoutSeconds and HeartbeatSeconds fields in the task state to trigger a retry.
AnswerB

The Retry field in a Step Functions task state allows you to specify retry behavior for failed tasks. Setting MaxAttempts to 3 and BackoffRate to 2.0 will retry the Glue job up to three times with exponential backoff (doubling the wait time between retries). This matches the requirement.

Why this answer

The Retry field in AWS Step Functions is specifically designed to handle retries on task failures, including automatic backoff. By specifying MaxAttempts and BackoffRate, the state machine will retry the Glue job invocation with exponential backoff, meeting the requirement without custom logic.

Exam trap

The trap here is confusing error handling with retry logic; Catch is for fallback, not for automatic retries with backoff.

325
MCQhard

A data pipeline ingests streaming data from Kinesis Data Streams into S3 via Kinesis Data Firehose. Occasionally, small files are written to S3, increasing downstream processing costs. What is the most efficient way to reduce the number of small files?

A.Use a Lambda function to aggregate records before sending to Firehose.
B.Use the Kinesis Client Library (KCL) to write larger batches to S3 directly.
C.Run a daily AWS Glue job to concatenate small files.
D.Increase the Firehose buffering interval to 300 seconds and buffering size to 64 MB.
AnswerD

Firehose buffers incoming records before delivering them, so raising the interval to 300 seconds and the size to 64 MB lets more records accumulate per delivery, producing fewer, larger S3 objects and cutting downstream processing overhead.

Why this answer

Kinesis Data Firehose allows you to configure buffering hints (size and interval) to control when data is delivered to S3. By increasing the buffering interval to 300 seconds and the buffering size to 64 MB, Firehose accumulates more records before writing, which reduces the number of small files. This is the most efficient approach as it requires no additional infrastructure or post-processing.

Exam trap

The trap here is that candidates may think a Lambda pre-aggregation (Option A) or a Glue job (Option C) is necessary, when in fact Firehose's built-in buffering configuration is the simplest and most cost-effective solution to control file sizes.

How to eliminate wrong answers

Option A is wrong because using a Lambda function to aggregate records before sending to Firehose adds latency and complexity, and Firehose already has built-in buffering capabilities that can be tuned without extra services. Option B is wrong because the Kinesis Client Library (KCL) is designed for consuming and processing records from a stream, not for writing directly to S3; it would require custom code to batch and write to S3, which is less efficient and not a managed solution. Option C is wrong because running a daily AWS Glue job to concatenate small files is a reactive, post-processing approach that does not prevent small files from being created in the first place, and it incurs additional compute costs and delays.

326
Multi-Selecthard

A data engineer is building a near-real-time ingestion pipeline into Amazon S3. Small JSON files arrive continuously from thousands of devices, and the engineer must optimize the data lake for downstream Amazon Athena queries while minimizing storage cost and query latency. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Store the raw JSON files in S3 Glacier Instant Retrieval to reduce storage cost immediately.
B.Configure an S3 Lifecycle rule to abort incomplete multipart uploads after 7 days.
C.Use AWS Glue ETL jobs to compact small JSON files into larger Parquet files partitioned by ingestion date.
D.Use Amazon Kinesis Data Firehose to buffer incoming records and deliver larger aggregated files to S3.
E.Enable S3 Transfer Acceleration on the ingestion bucket to speed up uploads from devices.
AnswersC, D

Compacting many small JSON files into larger Parquet files reduces the number of S3 objects and takes advantage of columnar storage, which lowers Athena scan costs and improves query latency. Partitioning by ingestion date further enables partition pruning. This directly addresses both the small-file problem and the need for efficient downstream analytics.

Why this answer

The pipeline suffers from many small JSON files, which increase Athena query overhead and cost. Compacting files into larger Parquet objects with AWS Glue ETL and using Kinesis Data Firehose to buffer and aggregate records both reduce object count and improve columnar query performance. Transfer Acceleration, Glacier Instant Retrieval, and aborting multipart uploads do not address file size or format for analytics.

Exam trap

The trap here is focusing on upload speed or storage class alone, when the real issue is the number and format of files that Athena must scan.

327
MCQhard

A company runs a data lake on AWS using S3 for storage and AWS Glue for ETL. The security team discovers that a contractor who left the company two months ago still has access to an S3 bucket containing sensitive data. The access was granted via an IAM user that was not deleted. The data engineer is asked to implement a solution to prevent future occurrences. The company uses AWS Organizations and has multiple accounts. The requirement is to automatically detect and remediate IAM users that have not been used for 90 days by disabling their access keys and notifying the security team. The solution must be least privilege and use AWS-native services. Which approach should the data engineer take?

A.Use AWS IAM Access Analyzer to generate findings for unused access and create an AWS Config managed rule to automatically disable the IAM user's access keys.
B.Use AWS CloudTrail to monitor IAM user activity and set up a CloudWatch alarm that triggers an SNS notification to the security team to manually disable the keys.
C.Use AWS Lake Formation to revoke the permissions of the IAM user and set up a scheduled Lambda function to check for unused IAM users.
D.Use AWS IAM Access Analyzer to generate findings for unused access and create an AWS Config custom rule with a Lambda function that automatically disables the access keys and sends a notification via SNS.
AnswerD

IAM Access Analyzer surfaces unused-access findings, and an AWS Config custom rule with Lambda disables stale access keys automatically, notifying via SNS. This satisfies the stem's least-privilege, AWS-native requirement for detecting and remediating 90-day-unused IAM users.

Why this answer

AWS IAM Access Analyzer can generate findings for unused access, and AWS Config with a custom rule can auto-remediate by invoking a Lambda function to disable keys and notify via SNS. Option A is incorrect because AWS Config managed rules cannot automatically disable access keys; they only evaluate compliance and require manual action or a custom rule with remediation. Option B is incorrect because CloudTrail and CloudWatch alarms do not automatically disable keys; they only notify for manual intervention.

Option C is incorrect because Lake Formation is used for fine-grained access control on data lakes, not for managing IAM users or disabling access keys.

328
MCQeasy

A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is stored as CSV files and is updated daily. The engineer wants to load only new data each day without duplicating existing records. Which AWS service or feature should the engineer use to automate this process?

A.Amazon Redshift Spectrum to query S3 data directly and insert new records.
B.Amazon Kinesis Data Firehose to stream S3 data to Redshift.
C.AWS Glue with job bookmarks enabled to track processed files.
D.AWS Database Migration Service (DMS) with ongoing replication from S3 to Redshift.
AnswerC

AWS Glue job bookmarks track the state of data processed in previous runs, allowing the job to process only new or changed files in subsequent runs. This is ideal for daily incremental loads from S3 to Redshift, as it prevents reprocessing and duplication. The engineer can create a Glue ETL job that reads from S3, applies transformations, and writes to Redshift using the COPY command. This automates the incremental load process efficiently.

Why this answer

AWS Glue job bookmarks are designed to track processed data across job runs, enabling incremental processing. For daily loads from S3 to Redshift, a Glue ETL job with bookmarks ensures only new files are processed each day, preventing duplication. This automates the process and integrates with Redshift via the COPY command.

Other options do not provide automatic incremental load tracking for S3 files.

Exam trap

The trap here is assuming that Redshift Spectrum or DMS can automatically handle incremental loads from S3 without additional logic.

329
MCQmedium

A data engineer must ingest a 4 TB Oracle database into Amazon S3 nightly. The database is on-premises and the network link supports only 200 Mbps. The engineer wants to minimize the total transfer time and avoid impacting production. Which approach should the engineer use?

A.Use Amazon S3 Transfer Acceleration with multipart uploads from an on-premises script.
B.Use AWS DataSync with an on-premises agent to transfer the data over the internet.
C.Use AWS Snowball Edge devices to transfer the data offline and import into Amazon S3.
D.Use AWS Database Migration Service (AWS DMS) with a replication instance and an S3 target endpoint.
AnswerC

AWS Snowball Edge is designed for large-scale offline data transfer. With 4 TB and a 200 Mbps link, online transfer would take roughly 44 hours or more. Snowball Edge can copy the data locally and ship it, avoiding the network bottleneck entirely. This minimizes transfer time and avoids saturating the production link.

Why this answer

For multi-terabyte datasets and limited bandwidth, offline transfer with AWS Snowball Edge is the recommended approach. It bypasses the network bottleneck, reduces transfer time, and avoids impacting production systems. Online methods like DMS, DataSync, or S3 Transfer Acceleration are constrained by the available bandwidth and are not optimal for this volume.

Exam trap

The trap here is assuming that an online acceleration service like S3 Transfer Acceleration or DataSync can overcome a slow origin uplink, when the bottleneck is the on-premises network itself.

330
MCQhard

An IAM role 'DataLakeRole' has the above S3 bucket policy attached to an S3 bucket. The role is assumed by an AWS Glue job. The Glue job is failing with 'Access Denied' errors when trying to list objects in the bucket. Which action should be added to the policy to fix the issue?

A.Add s3:ListObjects action for the bucket ARN.
B.Add s3:ListBucket action for the bucket ARN (arn:aws:s3:::my-data-lake).
C.Add s3:GetObjectVersion action for the object ARN.
D.Add s3:ListBucket action for the object ARN (arn:aws:s3:::my-data-lake/*).
AnswerB

Adding `s3:ListBucket` on the bucket ARN (`arn:aws:s3:::my-data-lake`) satisfies the failing list operation, because bucket-level actions such as `ListBucket` are evaluated against the bucket resource itself, not the objects within it. Object ARNs (`arn:aws:s3:::my-data-lake/*`) only govern object-level actions like `GetObject`.

Why this answer

The Glue job is failing with 'Access Denied' when trying to list objects, which requires the s3:ListBucket permission on the bucket itself (not on objects). Option B correctly adds s3:ListBucket for the bucket ARN (arn:aws:s3:::my-data-lake), which grants permission to list the contents of the bucket. Without this action, even if other permissions exist, the ListObjectsV2 API call used by AWS Glue to enumerate objects will be denied.

Exam trap

The trap here is that candidates confuse s3:ListBucket (which applies to the bucket itself) with s3:GetObject or s3:ListObjects (which are often misapplied to object ARNs), leading them to pick Option D or A, not realizing that listing requires the bucket-level permission and a bucket ARN, not an object ARN.

How to eliminate wrong answers

Option A is wrong because s3:ListObjects is an alias for s3:ListBucket and must be applied to the bucket ARN, not the bucket ARN with a trailing slash or object path; however, the key issue is that the action name itself is correct but the ARN in the answer is unspecified, and the question asks for the action to add, not the ARN—but more critically, s3:ListObjects is a legacy action name and the exam expects s3:ListBucket for consistency with the S3 API. Option C is wrong because s3:GetObjectVersion is used to retrieve a specific version of an object, not to list objects, and does not address the 'list objects' failure. Option D is wrong because s3:ListBucket must be applied to the bucket ARN (arn:aws:s3:::my-data-lake), not to an object ARN (arn:aws:s3:::my-data-lake/*); applying it to an object ARN would be invalid and would not grant the permission to list the bucket's contents.

331
Multi-Selecthard

A company stores sensitive financial data in an Amazon Redshift cluster. The data engineer must ensure that all queries are logged for audit purposes and that the logs are stored in Amazon S3 with server-side encryption. Which THREE steps should the data engineer take to meet these requirements?

Select 3 answers
A.Configure audit logs to be stored in an Amazon S3 bucket.
B.Enable encryption on the Redshift cluster.
C.Enable AWS CloudTrail to log Redshift queries.
D.Enable audit logging on the Redshift cluster.
E.Enable default encryption on the S3 bucket using SSE-S3 or SSE-KMS.
AnswersA, D, E

Audit logs can be delivered to an S3 bucket.

Why this answer

Amazon Redshift audit logs can be configured to be stored directly in an Amazon S3 bucket, which is a native feature for exporting connection logs, user logs, and query logs. This satisfies the requirement to log all queries for audit purposes without relying on external services.

Exam trap

The trap here is confusing AWS CloudTrail (which logs control-plane API calls) with Redshift's native audit logging (which logs data-plane SQL queries), leading candidates to incorrectly select CloudTrail as a solution for query auditing.

332
MCQmedium

Refer to the exhibit. A data engineer deploys this CloudFormation template to create an AWS Glue job. The job fails on the first run with an error: 'AccessDeniedException: User: arn:aws:sts::123456789012:assumed-role/GlueServiceRole/... is not authorized to perform: s3:GetObject on resource: s3://my-bucket/scripts/etl.py'. What is the most likely cause?

A.The ExecutionProperty MaxConcurrentRuns is set to 1, preventing the job from running.
B.The IAM role associated with the Glue job does not have an S3 GetObject permission for the script location.
C.The MaxRetries is set to 0, so the job does not retry on failure.
D.The script location is incorrectly specified; it should be an S3 URI with bucket and key.
AnswerB

AWS Glue assumes the job's IAM role to fetch the ETL script from S3 before execution begins. The AccessDeniedException naming s3:GetObject on the script path shows the role's policy lacks that permission, so the identity-based policy must grant GetObject on the script location.

Why this answer

The error message indicates that the IAM role 'GlueServiceRole' assumed by the AWS Glue job does not have the s3:GetObject permission for the script object at s3://my-bucket/scripts/etl.py. AWS Glue requires the execution role to have read access to the script location specified in the 'ScriptLocation' parameter. Without this permission, the job fails immediately on startup because it cannot download and execute the ETL script.

Exam trap

The DEA-C01 exam often tests the distinction between permissions errors and configuration errors, where candidates might incorrectly focus on script location format or job parameters instead of recognizing that an AccessDeniedException is a clear IAM permissions issue.

How to eliminate wrong answers

Option A is wrong because ExecutionProperty MaxConcurrentRuns controls how many concurrent runs of the job are allowed, not whether the job can start; it would not cause an AccessDeniedException. Option C is wrong because MaxRetries determines how many times the job retries after a failure, but the job fails on the first run with an access denied error, not a retry-related issue. Option D is wrong because the script location is already specified as an S3 URI (s3://my-bucket/scripts/etl.py), which is the correct format; the error is about permissions, not format.

333
MCQeasy

A data engineer is running an Amazon EMR cluster with Spark to process log files. The cluster uses instance fleets with m5.xlarge core nodes. The engineer observes that the Spark job is running slower than expected. CloudWatch metrics show that the cluster's CPU utilization is below 20% but memory utilization is near 90%. Which configuration change would most likely improve performance?

A.Use memory-optimized instances (r5.xlarge) for core nodes.
B.Increase the number of core nodes from 5 to 10.
C.Increase the number of Spark shuffle partitions.
D.Decrease the number of core nodes to reduce overhead.
AnswerA

Memory-optimised r5.xlarge instances provide a higher memory-to-vCPU ratio than m5.xlarge, directly relieving the near-90% memory utilisation that is throttling Spark executors. Since CPU sits below 20%, the bottleneck is memory capacity, not compute, so swapping core nodes to r5.xlarge lets executors hold larger partitions without spilling to disk.

Why this answer

The CloudWatch metrics show CPU below 20% while memory is near 90%, which is the classic signature of a memory-bound Spark workload. Spark executors on m5.xlarge (16 GiB RAM) are spilling to disk or GC-thrashing because the working set exceeds available heap. Switching core nodes to r5.xlarge (memory-optimized, 32 GiB RAM) doubles the memory per node, reducing spills and GC pressure, which directly addresses the bottleneck.

Exam trap

The trap here is assuming that 'slower than expected' always means insufficient compute, so candidates add nodes or partitions instead of reading the CloudWatch signal that memory — not CPU — is the saturated resource.

How to eliminate wrong answers

Option B is wrong because adding more core nodes increases aggregate memory but does not change the per-executor memory-to-core ratio; if each executor is already memory-starved, more nodes with the same instance type will still spill and the shuffle/IO overhead may even grow. Option C is wrong because increasing shuffle partitions addresses data skew or large partition sizes, not a memory bottleneck — it can actually increase overhead and small-file pressure. Option D is wrong because reducing core nodes lowers total cluster memory and parallelism, worsening the memory pressure and slowing the job further.

334
MCQhard

A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow and the system is experiencing high disk usage. The engineer suspects that the distribution style is suboptimal. Which action should the engineer take to improve query performance?

A.Convert all tables to use SORTKEY on the most frequently filtered column.
B.Increase the number of nodes in the cluster to distribute data across more slices.
C.Use the DISTSTYLE AUTO setting and analyze query patterns to let Redshift choose.
D.Set all tables to DISTSTYLE EVEN to distribute data evenly.
AnswerC

AUTO adapts distribution based on workload.

Why this answer

DISTSTYLE AUTO allows Amazon Redshift to automatically assign distribution styles (KEY, EVEN, or ALL) based on query patterns and table size, optimizing data distribution for improved query performance. This is particularly effective when the engineer suspects suboptimal distribution but lacks detailed knowledge of the ideal key, as Redshift analyzes workload patterns to reduce data movement and disk usage.

Exam trap

The trap here is that candidates often confuse distribution style with sort key or node scaling, leading them to choose options that address symptoms (e.g., disk usage via node count) rather than the root cause of suboptimal data distribution.

How to eliminate wrong answers

Option A is wrong because SORTKEY improves query performance by reducing the amount of data scanned via block-level filtering, but it does not address distribution style or high disk usage caused by data skew. Option B is wrong because increasing the number of nodes distributes data across more slices but does not fix the root cause of suboptimal distribution; it may even exacerbate disk usage if data is already skewed. Option D is wrong because setting all tables to DISTSTYLE EVEN distributes rows evenly across slices, which can eliminate data skew but may cause excessive data movement (broadcast or redistribution) during joins, degrading query performance for tables that are frequently joined on specific columns.

335
MCQmedium

A data engineer is using AWS Glue Studio to build a visual ETL job that joins a large Amazon S3 dataset with a small reference dataset of country codes. The join is currently implemented as a standard join, and the job runs slowly and shuffles large amounts of data. The engineer wants to optimize performance without changing the output. Which change should the engineer make?

A.Increase the job's timeout setting so the existing standard join has more time to complete.
B.Move the small reference dataset into the Glue Data Catalog as a view and reference it in the join.
C.Replace the standard join with a broadcast join so the small reference dataset is replicated to each executor.
D.Convert the large S3 dataset to a single file before the join to eliminate partitioning overhead.
AnswerC

When one side of a join is small, broadcasting it to every executor avoids shuffling the large dataset across the cluster. Glue Studio exposes a join type option that maps to Spark's broadcast join behavior. This reduces network and shuffle cost while producing the same joined output, directly addressing the slowness described.

Why this answer

A broadcast join replicates the small reference dataset to each executor, eliminating the shuffle of the large dataset that a standard join incurs. Because the output rows are unchanged, this is a pure performance optimization. Timeout increases, file consolidation, and catalog views do not alter the join execution strategy and therefore do not address the shuffle bottleneck.

Exam trap

The trap here is thinking a longer timeout fixes a slow join, when the real cost is shuffle volume that only a broadcast join can remove.

336
MCQhard

A company uses Kinesis Data Analytics for SQL-based real-time analytics on streaming data. They notice that the application is processing data slower than the incoming rate, causing increased latency. Which action is MOST likely to improve the throughput?

A.Increase the number of Kinesis Processing Units (KPUs) for the application
B.Increase the number of shards in the Kinesis data stream
C.Enable auto-scaling on the Kinesis data stream
D.Decrease the retention period of the Kinesis data stream
AnswerA

Kinesis Data Analytics parallelism scales with KPU count; each KPU supplies a fixed slice of CPU and memory, so raising KPUs lifts the application's processing ceiling above the incoming stream rate, directly addressing the throughput shortfall causing latency.

Why this answer

Kinesis Data Analytics for SQL applications processes data using Kinesis Processing Units (KPUs), which define the compute and memory resources available. When the incoming data rate exceeds the processing capacity, increasing the number of KPUs directly scales the application's parallelism and throughput, allowing it to keep up with the stream. This is the most direct way to reduce latency caused by insufficient processing power.

Exam trap

The trap here is that candidates often confuse scaling the source stream (shards) with scaling the analytics application (KPUs), assuming that more shards automatically improve processing throughput, when in fact the application's compute resources are the limiting factor.

How to eliminate wrong answers

Option B is wrong because increasing the number of shards in the Kinesis data stream increases the ingestion capacity and parallelism of the source stream, but it does not directly increase the processing capacity of the Kinesis Data Analytics application; the application must also be scaled (e.g., via KPUs) to consume the additional shards. Option C is wrong because enabling auto-scaling on the Kinesis data stream only adjusts the number of shards based on throughput, which again does not address the application's processing bottleneck. Option D is wrong because decreasing the retention period of the Kinesis data stream only reduces how long data is stored in the stream; it does not affect the processing rate or throughput of the analytics application.

337
MCQhard

A company uses AWS Glue to run ETL jobs that process data from Amazon RDS for MySQL and load it into Amazon S3. The job runs daily and processes incremental changes using the JDBC connection. Recently, the job has been failing with a 'Communications link failure' error. The RDS instance is in a private subnet. Which step should the engineer take first to diagnose the issue?

A.Verify that the IAM role used by Glue has the correct permissions to access RDS.
B.Change the Glue job type from Spark to Python shell.
C.Check the security group and network ACL rules for the RDS instance and the Glue connection.
D.Check that the JDBC driver is compatible with the Glue version.
AnswerC

A 'Communications link failure' means the JDBC connection cannot reach the private RDS endpoint. Verifying security group and network ACL rules for both the RDS instance and the Glue connection confirms whether the required port and subnet traffic are actually permitted.

Why this answer

The 'Communications link failure' error typically indicates a network connectivity issue between AWS Glue and the RDS instance. Since the RDS instance is in a private subnet, the Glue job must be able to reach it via a VPC endpoint or a Glue connection that uses network configuration. Checking the security group (inbound rules for the RDS instance allowing traffic from Glue's elastic network interfaces) and network ACLs (ensuring ephemeral ports are open) is the first logical step to diagnose connectivity.

Exam trap

The trap here is that candidates often jump to IAM permissions or JDBC driver issues first, but the 'Communications link failure' error is a classic network connectivity symptom that requires checking security groups and network ACLs before anything else.

How to eliminate wrong answers

Option A is wrong because IAM permissions control authentication and authorization to AWS services, not network-level connectivity; a 'Communications link failure' is a network error, not an access denied error. Option B is wrong because changing the job type from Spark to Python shell does not resolve network connectivity issues; it only changes the execution environment and may even introduce new limitations for JDBC connections. Option D is wrong because JDBC driver compatibility would cause a different error (e.g., 'No suitable driver' or class not found), not a 'Communications link failure', which is a network timeout or connection reset.

338
Multi-Selecthard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer needs to identify and resolve issues related to workload management (WLM). Which TWO actions should the engineer take to improve query performance? (Choose two.)

Select 2 answers
A.Increase the number of nodes in the cluster to add more compute resources.
B.Set the WLM query slots to the maximum value for all queues to increase concurrency.
C.Enable short query acceleration (SQA) to prioritize short-running queries.
D.Disable concurrency scaling to prevent additional clusters from being added.
E.Configure a manual WLM queue with a higher memory allocation for the queue handling complex queries.
AnswersC, E

Short query acceleration (SQA) uses machine learning to predict query execution time and runs short queries in a dedicated space, preventing them from waiting behind long-running queries. Enabling SQA can improve overall throughput and reduce latency for short queries during peak periods, making it an effective action to resolve WLM-related performance issues.

Why this answer

To improve query performance in Amazon Redshift during peak hours from a workload management perspective, the engineer should configure manual WLM queues with appropriate memory allocation and enable short query acceleration. Manual WLM allows fine-grained control over memory and concurrency, ensuring complex queries get necessary resources. SQA prioritizes short queries, reducing wait times.

These actions directly address WLM-related bottlenecks.

Exam trap

The trap here is confusing general scaling actions with WLM-specific tuning, such as adding nodes or maximizing concurrency, which can be counterproductive.

339
MCQhard

A data engineer is designing a multi-Region disaster recovery solution for an Amazon DynamoDB table. The table must be available in a secondary Region with minimal data loss and automatic failover. Which feature should be used?

A.DynamoDB on-demand backup and restore in the secondary Region
B.DynamoDB global tables
C.DynamoDB point-in-time recovery (PITR)
D.DynamoDB cross-Region snapshot export to S3
AnswerB

DynamoDB global tables replicate data across Regions with active-active writes and automatic failover, meeting the secondary-Region availability and minimal-data-loss requirements. Single-Region backups or point-in-time recovery cannot provide an operational secondary Region, so they fail the stated disaster recovery constraint.

Why this answer

DynamoDB global tables provide a fully managed, multi-Region, multi-active database solution that replicates data automatically across selected AWS Regions. This ensures automatic failover with eventual consistency and minimal data loss, meeting the disaster recovery requirements for high availability and automatic failover without manual intervention.

Exam trap

The trap here is that candidates often confuse point-in-time recovery (PITR) with cross-Region disaster recovery, but PITR is a single-Region feature that does not provide automatic failover or multi-Region replication.

How to eliminate wrong answers

Option A is wrong because on-demand backup and restore is a manual process that requires user intervention to initiate a restore in the secondary Region, not providing automatic failover or minimal data loss in real time. Option C is wrong because point-in-time recovery (PITR) protects against accidental writes or deletes within a single Region by restoring to a point in time, but it does not replicate data across Regions or enable automatic failover. Option D is wrong because cross-Region snapshot export to S3 is a manual, batch-oriented process that exports table data to Amazon S3 in another Region, requiring manual import and setup for failover, and does not provide automatic, continuous replication or failover.

340
Multi-Selectmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket encrypted with SSE-KMS and writes to another S3 bucket also encrypted with SSE-KMS. The job uses an IAM role. The security team requires that the job have only the minimum necessary permissions to decrypt and encrypt data. Which TWO actions should be included in the IAM policy attached to the Glue job role? (Choose two.)

Select 2 answers
A.kms:ReEncryptFrom on the KMS key used by the destination S3 bucket.
B.kms:Encrypt on the KMS key used by the source S3 bucket.
C.kms:CreateGrant on both KMS keys to allow S3 to use the keys.
D.kms:GenerateDataKey on the KMS key used by the destination S3 bucket.
E.kms:Decrypt on the KMS key used by the source S3 bucket.
AnswersD, E

When writing objects to the destination S3 bucket with SSE-KMS, S3 uses the KMS key to generate a data key for encryption. The Glue job role must have kms:GenerateDataKey permission on the destination key to allow S3 to encrypt the data on its behalf. This is required for the write operation to succeed.

Why this answer

To read SSE-KMS encrypted objects from the source bucket, the Glue job role needs kms:Decrypt on the source key. To write SSE-KMS encrypted objects to the destination bucket, S3 requires the job role to have kms:GenerateDataKey on the destination key. These two permissions are the minimum required for the job to function while adhering to least privilege.

Exam trap

The trap here is confusing the permissions needed for S3 to use KMS keys with permissions for direct KMS operations like CreateGrant or ReEncrypt.

341
Multi-Selectmedium

A company is building a data lake on Amazon S3. Data arrives from multiple sources in JSON, CSV, and Avro formats. The data must be transformed to Parquet and partitioned by date and source. Which TWO services can perform this transformation with minimal custom code? (Choose TWO.)

Select 2 answers
A.Amazon EMR with Spark
B.AWS Lake Formation
C.Amazon Athena CTAS queries
D.AWS Glue ETL jobs
E.Amazon Kinesis Data Firehose
AnswersA, D

Amazon EMR with Spark provides native readers for JSON, CSV and Avro plus a Parquet writer, and supports partitionBy on date and source columns. Transformations are expressed declaratively in Spark SQL or DataFrames, requiring little bespoke code.

Why this answer

Amazon EMR with Spark (A) is correct because Spark natively reads JSON, CSV, and Avro, and can write Parquet with partitionBy('date','source') using built-in DataFrame APIs, requiring only a small script rather than a custom transformation engine. AWS Glue ETL jobs (D) are correct because Glue provides managed, serverless Spark with built-in DynamicFrame readers/writers for JSON, CSV, and Avro, plus automatic schema inference and native Parquet output with partition keys, minimizing custom code. AWS Lake Formation (B) is a permission and metadata/catalog governance layer, not a data transformation engine.

Amazon Athena CTAS (C) can convert formats and partition results, but it is query-oriented and less suited to general multi-format ETL pipelines. Amazon Kinesis Data Firehose (E) is a streaming delivery service that can convert to Parquet via Glue schema, but it does not perform the required multi-source batch transformation and partitioning logic.

Exam trap

The trap here is that candidates often confuse AWS Lake Formation's data catalog and permission features with actual data transformation capabilities, or they assume Kinesis Data Firehose can transform existing S3 objects when it only processes streaming data in transit.

342
MCQmedium

A social media company ingests user activity data from multiple sources using Amazon Kinesis Data Firehose. The data is delivered to Amazon S3 in near-real-time. The company wants to transform the data by adding a timestamp and masking email addresses before storing it in S3. The transformation should be applied to all records. What is the most cost-effective way to implement this transformation?

A.Use Amazon Athena to run a CTAS query that transforms the data and writes to a new location.
B.Use AWS Glue to schedule a batch job every 5 minutes to transform the data.
C.Use Amazon S3 Events to trigger a Lambda function whenever a new object is created.
D.Configure the Firehose delivery stream to invoke a Lambda function for data transformation.
AnswerD

Firehose supports inline Lambda transformation, invoking the function on each batch before delivery to Amazon S3. This applies the timestamp addition and email masking to all records without managing servers, satisfying the stem's cost-effectiveness and universal-transformation constraints more cheaply than separate processing infrastructure.

Why this answer

Amazon Kinesis Data Firehose supports invoking an AWS Lambda function for inline data transformation before delivering data to the destination. This allows adding timestamps and masking email addresses in near-real-time as data flows through the delivery stream. It is the most cost-effective and operationally efficient way because it avoids additional storage, batch processing, or separate compute resources.

Exam trap

The trap is choosing batch or post-storage transformation methods (Glue, S3 Events) instead of inline transformation, which is more efficient and cost-effective for near-real-time streaming data.

How to eliminate wrong answers

Option A is wrong because Amazon Athena is a query service for analyzing data in S3, and running CTAS queries would require additional steps and is not designed for continuous near-real-time transformation; it would also incur costs for scanning data. Option B is wrong because AWS Glue batch jobs every 5 minutes introduce latency and require additional infrastructure and cost, and they are not integrated directly with Firehose. Option C is wrong because S3 Events triggering Lambda would process data after it is stored, adding latency and requiring additional Lambda invocations and possibly reprocessing, and it is not as seamless as Firehose transformation.

343
MCQeasy

A company needs to ensure that data stored in Amazon RDS is encrypted at rest. Which action should the data engineer take?

A.Enable encryption at rest by modifying the existing RDS instance.
B.Encrypt the underlying EBS volumes using AWS KMS.
C.Create a new RDS instance with encryption enabled using AWS KMS.
D.Enable SSL/TLS for connections to the RDS instance.
AnswerC

RDS encryption at rest can only be enabled when the instance is created; existing unencrypted instances cannot be encrypted in place. Creating a new instance with encryption enabled using AWS KMS therefore satisfies the requirement directly, unlike modifying an existing instance, which cannot add storage encryption.

Why this answer

Amazon RDS encryption at rest must be enabled when the DB instance is created, using AWS KMS. It cannot be added later. Option A is incorrect because encryption cannot be enabled on an existing RDS instance; you must create a new one with encryption enabled.

Option B is incorrect because encrypting the underlying EBS volumes does not encrypt the RDS database; RDS encryption at rest is separate and must be configured at the instance level. Option D is incorrect because SSL/TLS secures data in transit, not at rest.

344
MCQhard

A data engineer reviews the Glue job configuration. The job fails when processing large datasets. The error message indicates out-of-memory in the executors. Which change to the job configuration will most directly address this issue?

A.Change the worker type from Standard to G.2X.
B.Increase the timeout from 30 to 60 minutes.
C.Increase the number of workers from 5 to 10.
D.Set MaxRetries to 3.
AnswerA

G.2X workers have more memory (8 GB vs 4 GB), directly addressing OOM.

Why this answer

The job fails due to out-of-memory errors in the executors. The current configuration uses 5 Standard workers, each with 16 GB of memory. Changing the worker type to G.2X provides 32 GB per worker, doubling the memory per executor and directly addressing the OOM issue.

Increasing the number of workers (option C) adds more executors but does not increase memory per executor, which may not resolve the issue if each executor still runs out of memory. Options B and D do not affect memory.

345
MCQmedium

A company is using an Amazon RDS for MySQL database for its e-commerce platform. During a recent flash sale, the database experienced high read traffic, causing slow query performance. The company needs a solution that offloads read traffic with minimal application changes. Which action should be taken?

A.Enable DynamoDB Accelerator (DAX) on the RDS instance.
B.Migrate the database to Amazon Aurora and enable Aurora Global Database.
C.Implement Amazon ElastiCache for Redis to cache database queries.
D.Create an Amazon RDS read replica in the same region.
AnswerD

An RDS read replica uses MySQL's native asynchronous replication to serve read-only queries from a separate endpoint, diverting SELECT traffic from the primary. This offloads read load with minimal application change, satisfying the flash-sale read-traffic constraint.

Why this answer

Creating an Amazon RDS read replica in the same region offloads read traffic from the primary DB instance by directing read queries to a read-only copy. This requires minimal application changes—only modifying the database connection string to point read queries to the replica endpoint. RDS read replicas use MySQL's native asynchronous replication, making them ideal for scaling read-heavy workloads like flash sales.

Exam trap

The trap here is that candidates may choose ElastiCache (Option C) because it is a caching solution, but they overlook the explicit requirement for minimal application changes, which caching typically does not satisfy without code modifications.

How to eliminate wrong answers

Option A is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for Amazon DynamoDB, not for RDS for MySQL; it cannot be enabled on an RDS instance. Option B is wrong because migrating to Aurora and enabling Aurora Global Database is designed for cross-region disaster recovery and global reads, not for offloading read traffic within a single region, and it requires significant application and migration effort. Option C is wrong because while ElastiCache for Redis can cache query results, it requires application code changes to implement caching logic (e.g., cache-aside pattern), which contradicts the requirement for minimal application changes.

346
Multi-Selectmedium

Which TWO actions can help optimize Amazon S3 storage costs for a data lake? (Choose two.)

Select 2 answers
A.Enable S3 Replication to another region
B.Use S3 Intelligent-Tiering for unpredictable access patterns
C.Use S3 Select to retrieve only needed data
D.Enable S3 Transfer Acceleration
E.Implement S3 Lifecycle policies to transition objects to Glacier
AnswersB, E

S3 Intelligent-Tiering automatically moves objects between frequent and infrequent access tiers based on observed access patterns, charging a small monitoring fee. This satisfies the stem's unpredictable access constraint, since lifecycle policies require known patterns to schedule transitions reliably.

Why this answer

Option B is correct because S3 Intelligent-Tiering automatically moves objects between frequent and infrequent access tiers based on changing access patterns, eliminating retrieval fees and avoiding the cost of manual tiering when access is unpredictable. Option E is correct because S3 Lifecycle policies can transition aging objects to lower-cost storage classes such as S3 Glacier Instant Retrieval, Glacier Flexible Retrieval, or Glacier Deep Archive, directly reducing storage costs for cold data. Option A is not a cost optimization for a data lake because S3 Replication to another region incurs additional storage, request, and inter-region data transfer charges.

Option C reduces data scanned and compute cost for querying, but it does not lower S3 storage costs. Option D speeds up uploads over long distances via AWS edge locations but adds a premium charge, so it increases rather than optimizes storage cost.

Exam trap

The trap here is that candidates confuse cost optimization for storage (reducing stored data cost) with cost optimization for data transfer or retrieval, leading them to select options like S3 Select or Transfer Acceleration that address different cost dimensions.

347
MCQmedium

A company uses Amazon RDS for MySQL to store financial data. A compliance requirement mandates that all database connections must be encrypted. Which configuration step is necessary?

A.Set the RDS parameter require_secure_transport to 1.
B.Create the RDS DB instance in a private subnet.
C.Enable encryption for the RDS DB instance at creation time.
D.Configure the VPC security group to only allow traffic from certain IPs.
AnswerA

This is correct. To enforce encrypted connections for RDS MySQL, you must modify the DB parameter group to require SSL/TLS by setting parameters such as 'require_secure_transport' to 1. While the exact parameter name may vary, the intent is to enforce encryption in transit.

Why this answer

For Amazon RDS for MySQL, the parameter require_secure_transport controls whether the DB instance accepts only SSL/TLS-encrypted connections. Setting it to 1 enforces encryption for all client connections, satisfying the compliance mandate. This is the direct, database-level control for connection encryption, distinct from storage encryption or network isolation.

Exam trap

DEA-C01 often tests the confusion between encryption at rest (storage encryption) and encryption in transit (require_secure_transport), leading candidates to pick the storage encryption option for a connection-encryption requirement.

How to eliminate wrong answers

Option B is wrong because placing the DB instance in a private subnet restricts network reachability but does not encrypt connections — clients in the same VPC can still connect without TLS. Option C is wrong because enabling encryption at creation time encrypts data at rest (storage), not data in transit between clients and the database. Option D is wrong because security group rules filter source IPs and ports but do not enforce TLS on the connection itself.

348
MCQeasy

A company uses Amazon S3 to store customer documents. The data engineer needs to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted with a customer-managed AWS KMS key. What should the data engineer do?

A.Use pre-signed URLs for all uploads that include encryption parameters.
B.Create a bucket policy that denies uploads without encryption.
C.Enable S3 Versioning on the bucket.
D.Set default encryption on the bucket to use SSE-KMS with the customer-managed key.
AnswerD

Setting default bucket encryption to SSE-KMS with the customer-managed key enforces encryption at upload without changing client behaviour, satisfying the requirement that every object be encrypted automatically. S3 applies the default to objects lacking an explicit encryption header, and SSE-KMS provides the customer-managed key control the stem demands.

Why this answer

Setting default encryption on the S3 bucket to SSE-KMS with the customer-managed key ensures that all objects uploaded without explicit encryption headers are automatically encrypted using that KMS key. This satisfies the requirement without relying on client-side behavior, as S3 applies the encryption server-side at the time of write.

Exam trap

The trap here is that candidates often confuse bucket policies that deny unencrypted uploads (which only reject non-compliant requests) with default encryption (which automatically encrypts objects), leading them to choose Option B instead of D.

How to eliminate wrong answers

Option A is wrong because pre-signed URLs only grant temporary access to upload or download objects; they do not enforce encryption on the uploaded data, and including encryption parameters in the URL is optional and client-dependent. Option B is wrong because a bucket policy that denies uploads without encryption can enforce that clients must include encryption headers, but it does not automatically encrypt objects; if the client fails to include the header, the upload is denied rather than encrypted. Option C is wrong because S3 Versioning preserves multiple versions of an object but has no effect on encryption; it does not encrypt objects or enforce encryption policies.

349
Multi-Selecthard

A company is migrating its on-premises data warehouse to Amazon Redshift. The data includes tables with up to 100 columns and 500 million rows. The migration involves a full load followed by incremental updates. The company needs to minimize downtime during the final cutover. Which THREE strategies should the data engineer use to facilitate the migration? (Choose THREE.)

Select 3 answers
A.Increase the number of WLM queues to allow more concurrent loads.
B.Use the COPY command to load data from Amazon S3.
C.Use columnar format (e.g., Parquet) for the data files in S3.
D.Run VACUUM and ANALYZE commands after loading the data.
E.Disable distribution keys on the target tables to simplify loading.
AnswersB, C, D

The COPY command loads data in parallel from Amazon S3 into Redshift, directly satisfying the requirement to minimise cutover downtime during the full load. Its massively parallel architecture ingests the 500-million-row tables far faster than row-by-row INSERT statements, compressing the migration window before incremental updates begin.

Why this answer

Option B is correct because the COPY command is the most efficient, parallelized way to bulk-load large datasets from Amazon S3 into Redshift, and it supports loading from multiple files in parallel to speed up the full load and minimize cutover downtime. Option C is correct because storing the source files in a columnar format such as Parquet lets COPY read only needed columns, compresses data heavily, and reduces I/O and load time compared with row-based CSV or JSON. Option D is correct because after a large load, running VACUUM re-sorts rows and reclaims space (restoring sort-key performance) and ANALYZE refreshes table statistics so the query planner produces efficient plans for the subsequent incremental workload.

Option A is not appropriate because adding WLM queues does not by itself accelerate a single large load and can even reduce per-query resources; queue design is about workload isolation, not migration throughput. Option E is not appropriate because disabling distribution keys removes the ability to co-locate joins and causes data redistribution at query time, hurting performance; distribution keys should be chosen deliberately, not disabled.

Exam trap

DEA-C01 often tests whether candidates know that WLM queues and distribution keys are performance/concurrency constructs, not migration accelerators — the trap is picking 'more queues' or 'disable distribution keys' as shortcuts, when the real levers are COPY from S3, columnar formats, and post-load VACUUM/ANALYZE.

350
MCQmedium

A data engineer is designing a data pipeline that processes sensitive personal data. The data is ingested via Amazon Kinesis Data Firehose and stored in Amazon S3. The pipeline must ensure that the data is encrypted at rest and in transit. The engineer also needs to audit access to the data. Which combination of services meets these requirements?

A.AWS KMS for encryption at rest, Kinesis Data Analytics for in-transit encryption, and AWS CloudTrail for auditing.
B.AWS KMS for encryption at rest, Amazon CloudWatch Logs for auditing, and TLS for in-transit encryption.
C.S3 server-side encryption (SSE-S3) for at-rest encryption, HTTPS for in-transit encryption, and AWS CloudTrail for auditing.
D.S3 client-side encryption, AWS Config for auditing, and TLS for in-transit encryption.
AnswerC

SSE-S3 encrypts objects at rest, HTTPS secures data in transit from Firehose to S3, and CloudTrail records API activity for auditing. Together these three satisfy the encryption and access-audit requirements for the sensitive personal data pipeline.

Why this answer

Option C correctly combines S3 server-side encryption (SSE-S3) for data at rest, HTTPS (TLS) for data in transit, and AWS CloudTrail for auditing access. SSE-S3 provides AES-256 encryption managed by S3, HTTPS ensures secure ingestion and retrieval, and CloudTrail logs all API calls to S3, enabling audit trails. This meets all three requirements: encryption at rest, encryption in transit, and auditability.

Exam trap

DEA-C01 often tests the confusion between services that provide auditing (CloudTrail) versus monitoring (CloudWatch) or compliance (Config), and between encryption mechanisms for data at rest (SSE-S3, SSE-KMS) versus in transit (TLS/HTTPS).

How to eliminate wrong answers

Option A is wrong because Kinesis Data Analytics is an analytics service, not an encryption mechanism for in-transit data; in-transit encryption for Firehose is handled by HTTPS/TLS, not Kinesis Data Analytics. Option B is wrong because Amazon CloudWatch Logs is for monitoring and logging application/system metrics, not for auditing access to S3 data; auditing requires CloudTrail. Option D is wrong because AWS Config is a configuration compliance service, not an audit trail for data access; CloudTrail is the correct service for auditing access.

351
MCQmedium

A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format, and the engineer needs to convert it to Parquet before storage to optimize Athena queries. The Firehose delivery stream is configured with an AWS Lambda function for record transformation. However, the transformed data is still in JSON format in S3. What is the likely cause?

A.The Firehose delivery stream is not configured with the 'Convert record format' option to convert JSON to Parquet using an AWS Glue table.
B.The Lambda function is missing the necessary IAM permissions to write Parquet files to S3.
C.The Lambda function is not correctly converting the JSON to Parquet; it must output Parquet-formatted records.
D.The S3 bucket is configured with an S3 Lifecycle policy that converts objects to Parquet after a certain period.
AnswerA

Kinesis Data Firehose supports record format conversion from JSON to Parquet or ORC using an AWS Glue Data Catalog table. This feature must be explicitly enabled in the delivery stream settings. If it is not configured, the data remains in its original JSON format regardless of any Lambda transformation. The engineer needs to enable this option and specify the Glue table schema.

Why this answer

Kinesis Data Firehose can convert JSON to Parquet using the record format conversion feature, which relies on an AWS Glue table for schema. This option must be enabled in the delivery stream; otherwise, data remains in JSON even if a Lambda function transforms records. The Lambda function cannot perform the conversion itself, so enabling the Firehose setting is the correct fix.

Exam trap

The trap here is believing that a Lambda transformation function can output Parquet directly, when actually Firehose's built-in format conversion must be enabled.

352
Multi-Selecthard

A data engineer is configuring an AWS Glue crawler to catalog data in an Amazon S3 bucket. The bucket contains CSV files organized in folders by year and month, and new files are added daily. The engineer wants the crawler to detect schema changes automatically and avoid reprocessing unchanged files on subsequent runs. (Choose two.)

Select 2 answers
A.Enable the crawler's incremental crawling feature to identify and process only new folders and files since the last crawl.
B.Set the crawler's S3 target to exclude the folders that have already been cataloged using an exclude pattern.
C.Enable the crawler's schema change policy to update the table definition in the Data Catalog when new columns are detected.
D.Configure the crawler to use S3 event notifications so it runs only when new objects are created.
E.Change the crawler's output to create a separate table for each partition folder.
AnswersA, C

Incremental crawling lets the crawler track previously crawled S3 paths and process only new folders and files on subsequent runs. This directly satisfies the requirement to avoid reprocessing unchanged files, reducing crawl time and cost. Combined with a schema change policy, the crawler keeps the catalog current while minimizing redundant work.

Why this answer

Incremental crawling lets the crawler process only new folders and files since the last run, avoiding redundant scanning of unchanged data. The schema change policy, when set to update the table in the Data Catalog, automatically incorporates new columns detected during a crawl. Together these settings satisfy both requirements: detecting schema evolution and skipping already-processed files.

Exam trap

The trap here is assuming that S3 event notifications or exclude patterns will make the crawler incremental, when only the built-in incremental crawling feature tracks previously crawled paths and files.

353
Multi-Selectmedium

A data engineer is building an AWS Glue ETL job that reads large Parquet datasets from Amazon S3 and must optimize performance and cost. The engineer wants to reduce the number of small files written to the target S3 prefix and improve read efficiency. (Choose two.)

Select 2 answers
A.Set the Glue job's worker type to G.1X and increase the number of workers to the maximum allowed.
B.Convert the output format to CSV so that multiple small Parquet files can be merged by the S3 service automatically.
C.Use the AWS Glue groupFiles and groupSize options to coalesce small input files before processing.
D.Call coalesce or repartition on the DynamicFrame before writing to reduce the number of output files.
E.Enable the Glue job bookmark to skip previously processed files and reduce the number of files read.
AnswersC, D

groupFiles and groupSize allow the Glue reader to combine many small input files into larger groups, reducing the number of tasks and improving read efficiency. This directly addresses the small-file problem on the input side, which lowers overhead and improves job performance when reading Parquet datasets.

Why this answer

Input-side coalescing with groupFiles and groupSize reduces the number of read tasks, while output-side coalesce or repartition controls how many files are written. Together they attack the small-file problem from both ends, improving performance and lowering cost for Parquet datasets in S3.

Exam trap

The trap here is focusing on worker scaling or bookmarks for small-file issues, when the fix is input grouping and output partition control.

354
MCQhard

A data engineer is troubleshooting an Amazon Redshift cluster that is running out of disk space. The engineer runs STV_PARTITIONS and notices that some slices have significantly more data than others. What is the most likely cause and solution?

A.Poorly chosen sort keys; redefine sort keys
B.Data distribution skew due to uneven distribution style; change distribution style to EVEN or correct KEY
C.Some nodes are underutilized; add more nodes
D.Concurrency scaling is disabled; enable concurrency scaling
AnswerB

Uneven slice storage indicates distribution skew: rows cluster on certain slices under a poor KEY distribution style. Redistributing with EVEN, or choosing a higher-cardinality KEY column, spreads data evenly across slices and resolves the disk-space imbalance.

Why this answer

B is correct because STV_PARTITIONS shows per-slice disk usage, and significant variation indicates data distribution skew. Uneven distribution causes some slices to fill faster, leading to premature disk-full errors. Changing the distribution style to EVEN (for tables without join keys) or correcting the KEY distribution style (using a high-cardinality, evenly distributed column) rebalances data across slices.

Exam trap

The trap here is that candidates confuse sort keys (which improve query performance via zone maps) with distribution keys (which control data placement across slices), leading them to incorrectly select sort key redefinition as the fix for disk space skew.

How to eliminate wrong answers

Option A is wrong because sort keys affect query performance (min/max zone maps and block pruning), not how data is distributed across slices; disk space skew is a distribution issue, not a sort key issue. Option C is wrong because adding nodes increases total cluster capacity but does not fix existing data skew; the problem is uneven data placement, not insufficient total nodes. Option D is wrong because concurrency scaling handles workload bursts by adding transient compute capacity, not disk space; it does not affect how data is stored on existing slices.

355
MCQeasy

A data engineering team needs to transform CSV files to Parquet format after they land in an S3 bucket. The transformation should be triggered automatically as soon as a new file arrives. Which AWS service is best suited for this task?

A.AWS Batch job submitted by S3 event
B.Amazon EMR cluster running continuously
C.AWS Lambda function triggered by S3 event
D.AWS Glue ETL job scheduled every 5 minutes
AnswerC

S3 event notifications invoke a Lambda function within milliseconds of each object landing, satisfying the immediate-trigger requirement. Lambda natively transforms CSV to Parquet using libraries such as pandas with pyarrow, writing output back to S3, with no cluster provisioning or polling needed for this short, event-driven task.

Why this answer

AWS Lambda functions can be directly triggered by S3 events (e.g., `s3:ObjectCreated:*`) to process newly uploaded CSV files. This serverless approach provides near-instantaneous, event-driven transformation to Parquet without managing any infrastructure, making it the most cost-effective and simplest solution for this specific use case.

Exam trap

The trap here is that candidates often choose AWS Glue (Option D) because it is a dedicated ETL service, but they overlook the requirement for immediate, event-driven processing, which Glue's scheduled jobs cannot provide without additional event-bridge triggers.

How to eliminate wrong answers

Option A is wrong because AWS Batch requires provisioning compute resources and a job queue, adding latency and complexity for a simple file transformation that can be handled by a lightweight Lambda function. Option B is wrong because an Amazon EMR cluster running continuously incurs ongoing costs and management overhead, and is overkill for a simple CSV-to-Parquet conversion triggered by file arrival. Option D is wrong because a scheduled AWS Glue ETL job every 5 minutes introduces unnecessary polling and potential latency (up to 5 minutes), whereas the requirement is for immediate, event-driven processing.

356
MCQhard

A company stores data in Amazon S3 with server-side encryption using AWS KMS (SSE-KMS). The data engineer needs to give a third-party auditor read-only access to the encrypted objects. The auditor has an AWS account. Which strategy should be used?

A.Generate a presigned URL for each object the auditor needs to access.
B.Copy the objects to a new bucket encrypted with SSE-S3 and share that bucket.
C.Grant the auditor's IAM role permission to use the KMS key.
D.Update the S3 bucket policy to allow access from the auditor's account and update the KMS key policy to allow the auditor's account to decrypt.
AnswerD

SSE-KMS objects require both S3 authorisation and KMS decrypt permission, since S3 calls KMS on the reader's behalf. The bucket policy grants object access; the key policy grants kms:Decrypt to the auditor's account, enabling read-only access.

Why this answer

SSE-KMS requires two independent authorizations: the caller must have s3:GetObject on the bucket (via bucket policy or IAM) AND kms:Decrypt on the KMS key (via key policy or IAM). Because the auditor is in a separate AWS account, both the S3 bucket policy and the KMS key policy must explicitly grant the external account access. This is the only option that satisfies both layers for cross-account access.

Exam trap

DEA-C01 often tests the misconception that KMS key policy alone grants access to encrypted S3 objects, when in fact both the S3 bucket policy (for s3:GetObject) and the KMS key policy (for kms:Decrypt) must permit the cross-account principal.

How to eliminate wrong answers

Option A is wrong because presigned URLs are time-limited (max 7 days), impractical for auditing many objects, and still require the signer to have kms:Decrypt — they don't scale for auditor access patterns. Option B is wrong because copying to SSE-S3 removes the CMK requirement, violates the security posture, duplicates data unnecessarily, and SSE-S3 offers no key-level access control or audit trail. Option C is wrong because granting only KMS key permission is insufficient — without an S3 bucket policy or IAM permission for s3:GetObject, the auditor cannot even read the object, and cross-account access requires an explicit resource-based policy.

357
Multi-Selectmedium

Which TWO actions can improve the performance of an AWS Glue ETL job that processes large datasets in Amazon S3? (Choose two.)

Select 2 answers
A.Increase the frequency of the Glue crawler.
B.Use a single Availability Zone for the S3 bucket.
C.Increase the number of DPUs allocated to the job.
D.Use columnar file formats like Parquet or ORC.
E.Use a single large file instead of many small files.
AnswersC, D

Increasing DPUs adds parallel executors, so more partitions are processed concurrently during the shuffle and write stages. This directly addresses the stem's large-dataset constraint, where a single worker count throttles throughput. More DPUs raise aggregate memory and CPU, reducing spills to disk and shortening each stage's runtime.

Why this answer

Increasing the number of DPUs allocates more processing power to the Glue job, which can speed up data processing for large datasets. Option D is correct because columnar file formats like Parquet or ORC are more efficient for analytical queries, reduce I/O, and allow better compression compared to row-based formats. Option A is incorrect: increasing crawler frequency only affects the metadata catalog update frequency, not the ETL job performance.

Option B is incorrect: using a single Availability Zone for the S3 bucket does not improve performance and may reduce availability. Option E is incorrect: using a single large file can reduce parallelism, as distributed processing benefits from splitting data into multiple files to be processed in parallel by different executors.

358
MCQeasy

A data engineer needs to ensure that an Amazon Redshift cluster encrypts data at rest using a customer-managed AWS KMS key. Which configuration step is required?

A.Create a new cluster and select the default AWS managed key for encryption.
B.Create a new cluster and specify a customer-managed KMS key for encryption.
C.Use AWS CloudHSM to generate a key and attach it to the cluster.
D.Enable encryption on the existing cluster by modifying the cluster configuration.
AnswerB

A Redshift cluster's encryption key is fixed at creation time; you cannot swap the KMS key on an existing cluster. Creating a new cluster and specifying the customer-managed KMS key is therefore the required configuration step.

Why this answer

To encrypt an Amazon Redshift cluster with a customer-managed KMS key, you must specify that key at cluster creation time; Redshift does not allow you to change the encryption key of an existing cluster after it is created. The customer-managed key gives you control over key rotation, grants, and audit via CloudTrail, which is required when the default AWS-managed key (aws/redshift) does not satisfy compliance.

Exam trap

DEA-C01 often tests the misconception that you can modify an existing Redshift cluster to add or change encryption — in reality, encryption is fixed at creation and requires snapshot/restore to change.

How to eliminate wrong answers

Option A is wrong because the default AWS-managed key (aws/redshift) does not give the customer control over key policy, rotation, or grants, which is the stated requirement. Option C is wrong because AWS CloudHSM is a separate hardware security module service; Redshift integrates with KMS, not CloudHSM, for at-rest encryption keys. Option D is wrong because Redshift does not support enabling or changing encryption on an existing cluster in place — you must create a new encrypted cluster and migrate data (or restore from a snapshot with encryption).

359
Multi-Selecthard

A data engineer is configuring an AWS Glue ETL job that reads from and writes to an Amazon S3 bucket. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS, and that the job must fail if TLS is not used. Which two actions should the data engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Attach a bucket policy that denies s3:GetObject requests where aws:SecureTransport is false.
B.Enable default encryption on the S3 bucket using SSE-S3.
C.Configure the Glue job to use a VPC endpoint for Amazon S3 and set the endpoint policy to require TLS.
D.Attach a bucket policy that denies s3:PutObject requests where aws:SecureTransport is false.
E.Enable S3 Transfer Acceleration on the bucket.
AnswersA, D

A bucket policy denying s3:GetObject when aws:SecureTransport is false blocks any read request that does not use TLS. This ensures the Glue job cannot read data insecurely and will fail if it attempts a non-TLS connection, satisfying the requirement for encryption in transit on reads.

Why this answer

To enforce TLS for all data in transit between AWS Glue and Amazon S3, you must deny both read and write requests that do not use TLS. Bucket policies that deny s3:GetObject and s3:PutObject when aws:SecureTransport is false accomplish this. Default encryption, VPC endpoint policies, and Transfer Acceleration do not universally enforce TLS for all Glue-to-S3 traffic or cause the job to fail on insecure connections.

Exam trap

The trap here is assuming that enabling default encryption or using a VPC endpoint automatically enforces TLS for all traffic, when only explicit bucket policy denials based on aws:SecureTransport guarantee that insecure requests are rejected.

360
Multi-Selecteasy

A data engineer needs to audit data access in Amazon S3 for compliance. Which TWO services can be used to capture and analyze S3 access logs? (Choose TWO.)

Select 2 answers
A.Amazon CloudWatch Logs
B.AWS CloudTrail
C.S3 server access logs
D.Amazon Macie
E.AWS Config
AnswersB, C

AWS CloudTrail records S3 data-plane API calls such as GetObject and PutObject, capturing the identity, source IP and timestamp of each access. This satisfies the audit requirement by providing an analysable event trail for compliance, complementing S3 server access logs.

Why this answer

AWS CloudTrail (B) is correct because it records S3 data events (object-level operations like GetObject and PutObject) and management events, delivering them to CloudTrail logs or CloudWatch Logs for auditing and analysis of who accessed what in S3. S3 server access logs (C) are correct because they provide detailed, best-effort records of every request made to a bucket, including requester, bucket, key, operation, and response codes, which can be stored and analyzed for compliance auditing. Amazon CloudWatch Logs (A) is not a log source for S3 access itself; it only stores and analyzes logs that are delivered to it by other services such as CloudTrail.

Amazon Macie (D) is a data security service that discovers and classifies sensitive data in S3, not a service for capturing S3 access logs. AWS Config (E) tracks resource configuration changes and compliance against rules, not individual S3 data access events.

361
MCQeasy

A company uses Amazon DynamoDB to store user session data. The table has a partition key of UserID and a sort key of SessionStartTime. The application frequently queries for all sessions of a specific user within a date range. The table is provisioned with 1000 RCUs and 1000 WCUs. During peak hours, the application experiences throttling on read requests. Which action should a data engineer take to resolve the throttling with minimal changes?

A.Increase the provisioned read capacity units (RCUs) for the table.
B.Create a global secondary index (GSI) on the SessionStartTime attribute.
C.Enable DynamoDB Accelerator (DAX) for the table.
D.Change the table's capacity mode to on-demand.
AnswerA

Throttling on read requests indicates insufficient read capacity. Increasing RCUs directly addresses the issue by allowing more read operations per second. Since the table is already provisioned, adjusting RCUs is a minimal change that can be done without application modifications or schema changes.

Why this answer

Read throttling occurs when the consumed read capacity exceeds the provisioned RCUs. Increasing RCUs is the most direct and minimal solution, as it provides more read capacity without requiring application changes or architectural modifications. It addresses the root cause of the throttling.

Exam trap

The trap here is assuming that a more complex solution like DAX or on-demand mode is needed, when simply increasing provisioned capacity resolves the immediate throttling issue.

362
Multi-Selectmedium

A data engineer is preparing an AWS Glue ETL job that reads from and writes to Amazon S3 and must audit every access to sensitive data for compliance. The security team wants to know which principals accessed which objects and when, and also wants to detect anomalous access patterns. Which TWO AWS services should be used together to meet these requirements? (Choose two.)

Select 2 answers
A.Amazon Macie sensitive data discovery jobs
B.Amazon S3 Inventory reports
C.AWS Glue job bookmarks
D.Amazon GuardDuty S3 Protection
E.AWS CloudTrail data events for S3
AnswersD, E

GuardDuty S3 Protection continuously analyzes CloudTrail data and management events to identify anomalous or suspicious S3 access, such as unusual API calls or access from unexpected locations. It surfaces findings that help the security team detect behavior changes rather than only recording raw events. Used alongside CloudTrail data events, it covers both the audit record and the anomaly detection requirement.

Why this answer

CloudTrail data events capture object-level API activity with caller identity, bucket, key, and time, forming the required audit trail for sensitive object access. GuardDuty S3 Protection analyzes that activity to surface anomalous or suspicious access. Together they provide both the detailed record and the detection capability, while inventory, bookmarks, and Macie address classification or processing state rather than access auditing.

Exam trap

The trap here is confusing data classification and inventory services such as Macie or S3 Inventory with access auditing, which requires CloudTrail data events.

363
MCQmedium

A data engineer manages a large Amazon S3 data lake with millions of small JSON files ingested daily. Amazon Athena queries against this data lake are slow and costly due to high per-query data scanned. The engineer wants to optimize the storage layout to improve query performance and reduce cost, while keeping the data queryable in place. Which solution should the engineer implement?

A.Use AWS Glue ETL to compact the small JSON files into larger Parquet files partitioned by commonly filtered columns, then update the AWS Glue Data Catalog.
B.Enable S3 Transfer Acceleration on the bucket to speed up data retrieval by Athena.
C.Convert the data to CSV format and add more columns to the AWS Glue Data Catalog table.
D.Move the data to Amazon Redshift and query it using Redshift Spectrum.
AnswerA

Compacting small files into larger Parquet files reduces the number of S3 GET requests and leverages columnar storage, which minimizes data scanned by Athena. Partitioning by frequently filtered columns enables partition pruning, further reducing data scanned. Updating the Data Catalog ensures Athena queries use the new schema and partitions. This directly addresses performance and cost without moving data out of S3.

Why this answer

Compacting small files into larger Parquet files and partitioning by frequently filtered columns directly reduces the amount of data scanned by Athena, improving performance and lowering cost. Parquet's columnar format allows Athena to read only the columns needed, and partitioning enables partition pruning. Updating the Data Catalog ensures queries use the optimized layout.

The other options do not address the core issues of small files and inefficient format.

Exam trap

The trap here is assuming that S3 Transfer Acceleration or format changes alone improve Athena query performance, when the key is reducing data scanned through compaction and partitioning.

364
Multi-Selectmedium

A data engineer is configuring an AWS Glue crawler to catalog CSV files stored in Amazon S3. The files are organized in prefixes by year and month, and the engineer wants the crawler to detect new partitions automatically and avoid re-crawling unchanged partitions. (Choose two.)

Select 2 answers
A.Set the crawler's schema change policy to update the table definition in the Data Catalog.
B.Configure the crawler to use incremental crawling so it processes only folders added since the last crawl.
C.Ensure the S3 prefixes follow a Hive-style partition naming convention with key=value pairs.
D.Add a path in the crawler configuration that points to each year and month prefix individually.
E.Create a partition index on the Data Catalog table to speed up partition filtering.
AnswersB, C

Incremental crawling lets a Glue crawler examine only new partitions or folders that appeared since the previous run, rather than rescanning the entire S3 prefix. This directly reduces crawl time and cost for a partitioned layout organized by year and month, satisfying the requirement to detect new partitions automatically while avoiding re-crawling unchanged partitions.

Why this answer

Incremental crawling limits each crawler run to newly added folders, so previously cataloged partitions are not rescanned, reducing time and cost. Hive-style key=value prefixes make the year and month directories recognizable as partitions so the crawler can populate them correctly in the Data Catalog. Together they deliver automatic partition detection without reprocessing unchanged data.

Exam trap

The trap here is believing that schema change policies or partition indexes control crawler partition discovery, when incremental crawling and Hive-style layout do.

365
Multi-Selecthard

A healthcare company stores patient records in an Amazon S3 bucket encrypted with SSE-KMS using a customer managed key. A new AWS Glue ETL job must read these records and write transformed data to another S3 bucket that is also encrypted with the same KMS key. The company's security policy requires that the Glue job's access to the KMS key be least-privilege and auditable. Which TWO actions should the data engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable S3 Block Public Access on both S3 buckets and enable default encryption with SSE-S3 to simplify key management.
B.Grant the Glue service role kms:* permissions on all KMS keys in the account to ensure the job can read and write data without interruption.
C.Attach a key policy to the KMS key that allows the AWS Glue service role to use the key for decrypt and generateDataKey operations, scoped to the specific S3 buckets via encryption context conditions.
D.Use AWS Lake Formation to grant the Glue job fine-grained access to the underlying S3 data and rely on Lake Formation to manage KMS permissions automatically.
E.Create an IAM policy that allows the Glue service role to call kms:Decrypt and kms:GenerateDataKey on the specific KMS key ARN, and attach it to the role.
AnswersC, E

The key policy is the primary access control for a KMS key. Granting the Glue service role only kms:Decrypt and kms:GenerateDataKey, with encryption context conditions that match the bucket ARNs, enforces least privilege and ensures the key can only be used for the intended S3 data. This directly satisfies the auditable, least-privilege requirement for the Glue job.

Why this answer

Access to SSE-KMS encrypted S3 objects requires permissions in both the KMS key policy and the IAM identity policy of the calling principal. The Glue service role must be granted only the necessary KMS operations (Decrypt and GenerateDataKey) on the specific key, with conditions to scope usage. This combination enforces least privilege and provides an auditable trail.

Broad permissions or unrelated services do not meet the security policy.

Exam trap

The trap here is assuming that granting IAM permissions alone or using Lake Formation is sufficient for a Glue job to read SSE-KMS encrypted S3 data, when the KMS key policy must also explicitly allow the role.

366
MCQmedium

A data engineer needs to run an AWS Glue extract, transform, and load (ETL) job that joins an Amazon S3-based Parquet dataset with a slowly changing dimension table in Amazon Redshift. The Redshift cluster is in a private subnet and cannot be reached over the public internet. The engineer wants the Glue job to read from Redshift without exposing credentials in the job script. Which combination of actions should the engineer take to meet these requirements?

A.Copy the Redshift data to Amazon S3 using an UNLOAD command, then have the Glue job read the S3 copy, and delete the S3 copy after the job completes.
B.Attach the Glue job to a VPC connection with a NAT gateway, store the Redshift credentials in AWS Secrets Manager, and reference the secret in the Glue job's connection options.
C.Use the Redshift Data API from the Glue job, attaching an IAM role to the job that has redshift-data:ExecuteStatement permissions, and run all transformations in Redshift.
D.Create an AWS Glue connection of type JDBC to the Redshift cluster, attach the job to a VPC connection in the same VPC, and store the Redshift credentials in AWS Secrets Manager for the connection to use.
AnswerD

A JDBC Glue connection combined with a VPC connection places the Glue job's elastic network interfaces in the same VPC as Redshift, enabling private connectivity. Storing credentials in Secrets Manager lets the connection retrieve them at runtime without embedding them in the script, satisfying both the network and credential requirements.

Why this answer

A JDBC connection plus a VPC connection gives the Glue job a private network path to the Redshift cluster, which is required because the cluster is in a private subnet. Storing credentials in Secrets Manager and referencing them through the connection avoids hardcoding secrets in the job script. Together these satisfy both the private connectivity and credential management requirements.

Exam trap

The trap here is assuming that adding a NAT gateway to a VPC connection is enough to reach a private Redshift cluster, when the job actually needs network interfaces in the same VPC.

367
MCQmedium

A data engineer must ensure that an AWS Glue ETL job can read from an Amazon S3 bucket encrypted with SSE-KMS and write to another S3 bucket also encrypted with SSE-KMS, using a single KMS key. The engineer has created an IAM role for the Glue job with permissions to access both buckets. What additional step is required to allow the Glue job to decrypt and encrypt data using the KMS key?

A.Attach a KMS key policy to the customer managed key that allows the Glue job's IAM role to use the key for encrypt and decrypt operations.
B.Grant the Glue job's IAM role the kms:CreateGrant permission in its identity-based policy.
C.Enable default encryption on both S3 buckets using SSE-S3 instead of SSE-KMS.
D.Modify the S3 bucket policy to allow the Glue job's IAM role to perform kms:Decrypt on the bucket.
AnswerA

KMS key policies are the primary way to control access to customer managed keys. Even if the IAM role has permissions in its identity-based policy, the key policy must also grant access. For a Glue job to use a KMS key, the key policy must explicitly allow the IAM role to perform kms:Decrypt and kms:Encrypt. Without this, the job will fail with access denied.

Why this answer

For a Glue job to use a KMS key for SSE-KMS encryption, both the IAM role's identity-based policy and the KMS key policy must grant the necessary permissions. The key policy is the resource-based policy that controls access to the KMS key. Without an explicit allow in the key policy for the Glue job's IAM role to perform kms:Decrypt and kms:Encrypt, the job cannot use the key, even if the IAM policy allows it.

Exam trap

The trap here is assuming that IAM permissions alone are sufficient to use a KMS key, when the KMS key policy must also grant access.

368
MCQmedium

A company stores sensitive data in an S3 bucket. To meet compliance requirements, they must ensure that all objects are encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement should be applied to deny uploads that do not use the required encryption?

A.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption":"aws:kms"}}}
B.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption":"AES256"}}}
C.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption-aws-kms-key-id":"arn:aws:kms:us-east-1:123456789012:key/abc123"}}}
D.{"Effect":"Deny","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucketname/*","Condition":{"Null":{"s3:x-amz-server-side-encryption":"true"}}}
AnswerA

The `StringNotEquals` condition on `s3:x-amz-server-side-encryption` denies any `PutObject` request whose encryption header is not `aws:kms`, directly enforcing the compliance requirement that all objects use SSE-KMS. Unlike `StringNotEqualsIfExists`, this operator also blocks uploads that omit the header entirely, closing the unencrypted-upload gap.

Why this answer

It uses the `s3:x-amz-server-side-encryption` condition key with `StringNotEquals` set to `aws:kms`, which denies any `s3:PutObject` request where the encryption header does not specify `aws:kms`. This ensures all uploaded objects are encrypted at rest using server-side encryption with AWS KMS, meeting the compliance requirement.

Exam trap

The trap here is that candidates often confuse the condition keys for encryption type (`s3:x-amz-server-side-encryption`) with the specific KMS key ID (`s3:x-amz-server-side-encryption-aws-kms-key-id`), or mistakenly use `Null` to check for the presence of encryption instead of enforcing the correct encryption algorithm.

How to eliminate wrong answers

Option B is wrong because it checks for `AES256`, which corresponds to SSE-S3 (Amazon S3-managed keys), not SSE-KMS (AWS KMS keys), so it would allow objects encrypted with SSE-S3 instead of enforcing KMS encryption. Option C is wrong because it uses the `s3:x-amz-server-side-encryption-aws-kms-key-id` condition key to require a specific KMS key ID, but the requirement is only to use KMS encryption, not a particular key; this would deny uploads using any other KMS key, even if they use KMS encryption. Option D is wrong because it uses the `Null` condition to deny uploads where the `s3:x-amz-server-side-encryption` header is not present (i.e., null), but it does not enforce that the encryption must be `aws:kms`; it would also allow uploads with `AES256` or other encryption values, failing to meet the KMS-specific requirement.

369
MCQmedium

A data engineer manages an AWS Glue job that processes JSON files from Amazon S3 and writes Parquet to another S3 location. The job intermittently fails with 'Unable to find catalog table' errors, even though the table exists in the Glue Data Catalog. The job's IAM role has full S3 access but only limited Glue permissions. Which action will resolve the failure with the LEAST privilege?

A.Enable AWS Glue Data Catalog encryption and update the job's security configuration.
B.Add glue:GetTable and glue:GetDatabase permissions for the specific database and table to the job's IAM role.
C.Attach the AWSGlueServiceRole managed policy to the job's IAM role.
D.Modify the job to use a different IAM role that has s3:GetObject permissions on the catalog bucket.
AnswerB

The error indicates the job cannot retrieve table metadata from the Data Catalog. Granting glue:GetTable and glue:GetDatabase scoped to the exact database and table provides the necessary read access without excessive permissions. This aligns with least privilege and directly addresses the missing permission that causes the catalog lookup to fail during job execution.

Why this answer

The job fails because its IAM role cannot call Glue Data Catalog APIs to retrieve table metadata. The minimal fix is to grant glue:GetTable and glue:GetDatabase for the specific database and table. This resolves the error while adhering to least privilege, unlike broad managed policies or irrelevant S3 permissions.

Exam trap

The trap here is assuming that S3 permissions alone are sufficient for AWS Glue jobs, when in fact the Glue Data Catalog requires its own IAM permissions.

370
Multi-Selectmedium

A company is using Amazon S3 to store sensitive data. They need to ensure that all objects are encrypted at rest. Which combination of actions should be taken? (Choose TWO.)

Select 2 answers
A.Enable S3 Versioning on the bucket.
B.Enable MFA Delete on the bucket.
C.Configure S3 Access Points with network policies.
D.Use a bucket policy to deny PutObject requests that do not include the x-amz-server-side-encryption header.
E.Enable default encryption on the S3 bucket.
AnswersD, E

The bucket policy enforces encryption at upload time by rejecting any PutObject lacking the x-amz-server-side-encryption header, so unencrypted writes cannot succeed. This satisfies the requirement that all objects be encrypted at rest by preventing plaintext objects from ever being stored.

Why this answer

A bucket policy that denies PutObject requests lacking the `x-amz-server-side-encryption` header enforces encryption at the time of upload, ensuring that any object written without explicit encryption headers is rejected. Option E is correct because enabling default encryption on the S3 bucket automatically applies server-side encryption (SSE-S3 or SSE-KMS) to any object uploaded without specifying encryption headers, providing a fallback that covers all objects. Together, these actions ensure that every object stored in the bucket is encrypted at rest, either by explicit client request or by default bucket settings.

Exam trap

The trap here is that candidates often confuse data protection features like Versioning or MFA Delete with encryption controls, or assume that network policies (Access Points) somehow enforce encryption, when in reality only explicit bucket policies and default encryption settings directly ensure objects are encrypted at rest.

371
MCQeasy

A data engineer needs to store semi-structured JSON logs from multiple sources in a centralized data store for querying using SQL. The logs are immutable and need to be retained for 90 days. Which AWS service should be used?

A.Amazon RDS for MySQL.
B.Amazon DynamoDB.
C.Amazon S3 with Amazon Athena.
D.Amazon ElastiCache for Redis.
AnswerC

Amazon S3 stores immutable JSON objects durably at low cost, while Athena queries them in place using standard SQL through the Glue Data Catalog. This satisfies the semi-structured, SQL-queryable, 90-day retention constraints without loading data into a separate database.

Why this answer

Amazon S3 with Amazon Athena is the correct choice because S3 provides durable, cost-effective storage for immutable semi-structured JSON logs, and Athena enables serverless SQL querying directly against the data in S3 without needing to load or transform it. This combination meets the 90-day retention requirement and supports querying semi-structured data using standard SQL via Athena's built-in JSON SerDe.

Exam trap

The trap here is that candidates may choose DynamoDB for its JSON support and querying flexibility, overlooking that it is not designed for cost-effective long-term retention of immutable logs and lacks native SQL querying, while S3 with Athena directly addresses both requirements.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for MySQL is a relational database designed for structured data with predefined schemas, not optimized for storing large volumes of immutable semi-structured JSON logs, and it incurs higher costs for long-term retention. Option B is wrong because Amazon DynamoDB is a NoSQL key-value and document database that can store JSON, but it is not cost-effective for 90-day retention of immutable logs due to per-request pricing and storage costs, and it lacks native SQL querying capabilities without additional services like DynamoDB Accelerator or PartiQL. Option D is wrong because Amazon ElastiCache for Redis is an in-memory cache designed for low-latency access to transient data, not for durable, long-term storage of immutable logs, and it does not support SQL querying.

372
Multi-Selectmedium

A company is designing a data lake on Amazon S3. The security team requires granular access control based on data classifications. Which TWO AWS services can be used together to implement attribute-based access control (ABAC) for objects in S3?

Select 2 answers
A.AWS Secrets Manager
B.AWS Lake Formation
C.Amazon S3 object tags
D.AWS Identity and Access Management (IAM)
E.AWS Key Management Service (KMS)
AnswersC, D

S3 object tags supply the resource-side attributes that ABAC evaluates, letting bucket policies and IAM permissions grant or deny access based on classification tags such as data sensitivity. Combined with IAM, tags form the attribute pair required for tag-based, granular object access control.

Why this answer

Amazon S3 object tags (C) are the attribute source in S3 ABAC: classification labels such as DataClass=Confidential are attached to objects, and IAM policies reference them via conditions like s3:ExistingObjectTag/<key>. AWS Identity and Access Management (D) is the policy engine that evaluates those tag-based conditions in identity and resource policies, granting or denying access according to the object's classification attribute. Together, S3 object tags supply the attributes and IAM enforces attribute-based access control, which is exactly the ABAC pattern required for the data lake.

AWS Secrets Manager (A) stores and rotates secrets such as credentials, not access-control attributes. AWS Lake Formation (B) manages fine-grained permissions over data catalog tables and databases, not S3 object-level ABAC. AWS Key Management Service (E) provides encryption key management and cryptographic operations, not authorization decisions.

373
MCQmedium

A company has an Amazon S3 bucket with versioning enabled. They want to automatically delete noncurrent versions of objects after 30 days. Which lifecycle rule action should be used?

A.Expiration
B.NoncurrentVersionExpiration
C.NoncurrentVersionTransition
D.AbortIncompleteMultipartUpload
AnswerB

NoncurrentVersionExpiration targets noncurrent object versions specifically, permanently deleting them once the specified retention period elapses. This satisfies the stem's 30-day constraint on noncurrent versions, unlike CurrentVersionExpiration, which acts only on the active version. Versioning must be enabled, which the scenario confirms.

Why this answer

The NoncurrentVersionExpiration lifecycle action is specifically designed to remove noncurrent object versions after a specified number of days. Since versioning is enabled and the requirement is to delete noncurrent versions after 30 days, this action directly meets the goal without affecting current versions or other lifecycle aspects.

Exam trap

The trap here is confusing NoncurrentVersionExpiration with Expiration, as candidates often mistakenly apply the standard Expiration action to delete old versions, not realizing it only affects the current version.

How to eliminate wrong answers

Option A is wrong because Expiration deletes the current version of an object (or marks it for deletion in non-versioned buckets), not noncurrent versions. Option C is wrong because NoncurrentVersionTransition moves noncurrent versions to a different storage class (e.g., S3 Glacier), but does not delete them. Option D is wrong because AbortIncompleteMultipartUpload only aborts incomplete multipart uploads that are older than a specified number of days, and has no effect on existing object versions.

374
MCQhard

An e-commerce company uses AWS Glue to run ETL jobs that transform clickstream data from Amazon S3. The job reads Parquet files, performs aggregations, and writes the results to Amazon Redshift. The job runs successfully but takes longer than expected. The data volume is increasing. Which design change would MOST improve the job's performance?

A.Write the aggregated results to a single large file instead of multiple partitions.
B.Convert the Parquet files to CSV to simplify the schema.
C.Replace the Redshift target with Amazon Redshift Spectrum.
D.Increase the number of Glue worker nodes (DPUs) for the job.
AnswerD

Adding DPUs increases the number of executors available to process partitions in parallel, directly addressing the growing data volume that is stretching job duration. Glue scales horizontally, so more workers reduce per-node workload and shorten the aggregation and write phases.

Why this answer

Increasing the number of Glue worker nodes (DPUs) directly scales the distributed processing capacity of the ETL job, allowing it to process larger volumes of Parquet data in parallel. This is the most straightforward way to reduce execution time when data volume is growing, as AWS Glue automatically partitions the workload across the additional workers.

Exam trap

The trap here is that candidates assume increasing DPUs always increases cost without considering that the job's runtime reduction often lowers total cost, and they mistakenly choose a data format or target change that does not address the core parallelism issue.

How to eliminate wrong answers

Option A is wrong because writing to a single large file eliminates parallelism in downstream reads and can cause bottlenecks in Redshift's COPY operation, which benefits from multiple files for concurrent loading. Option B is wrong because converting Parquet to CSV increases file size and I/O overhead due to lack of columnar compression and predicate pushdown, degrading performance. Option C is wrong because replacing Redshift with Redshift Spectrum would offload query processing to S3 but does not address the ETL job's performance bottleneck; the job still writes to Redshift, and Spectrum is a query engine, not a write target.

375
MCQmedium

An Amazon Kinesis Data Streams application is lagging behind. The data records are small (1 KB) and the shard count is 10. The consumer uses the KCL with default configuration. Which action will MOST effectively reduce the consumer lag?

A.Increase the number of KCL workers per shard (e.g., 2 workers per shard).
B.Use Enhanced Fan-Out to provide dedicated throughput.
C.Increase the number of shards to 20.
D.Reduce the record size by compressing the data.
AnswerB

Enhanced Fan-Out provides each consumer with dedicated throughput (2 MB/s per shard) and push-based delivery, which reduces latency and lag directly. This is the most effective action to reduce consumer lag given the scenario.

Why this answer

Enhanced Fan-Out provides each consumer with dedicated throughput (2 MB/s per shard) and eliminates the need for polling, which reduces latency and lag. Given that the consumer is lagging with default configuration and small records, Enhanced Fan-Out directly addresses the consumer-side bottleneck by providing dedicated read throughput, making it the most effective option. Option A is incorrect because the KCL does not support multiple workers per shard; each shard is processed by a single worker in a single-threaded manner.

Option C may help if the shard is saturated with incoming data, but the problem is consumer lag, not data ingestion. Option D reduces data size but does not address the consumer processing speed.

Exam trap

Candidates often assume that adding more shards (Option C) will always reduce lag, but the bottleneck is the consumer's processing capacity. Enhanced Fan-Out (Option B) is specifically designed to reduce consumer lag by providing dedicated throughput per consumer.

How to eliminate wrong answers

Option B is wrong because Enhanced Fan-Out provides dedicated throughput per consumer (up to 2 MB/s per shard per consumer), but the issue here is processing lag, not throttling or throughput limits—the default KCL already handles the 1 KB records easily, so dedicated throughput does not address the processing bottleneck. Option C is wrong because increasing shards to 20 would increase the number of parallel processing units, but each shard still has only one KCL worker by default, so the per-shard processing capacity remains unchanged; this would only help if the shard were overloaded with data, which is not the case with small records. Option D is wrong because compressing data reduces the size of records, but the records are already only 1 KB, and the bottleneck is processing time per record, not network or storage throughput; compression adds CPU overhead and does not reduce lag.

Page 4

Page 5 of 18

Page 6