DEA-C01 · domain
troubleshooting
Practise AWS Certified Data Engineer Associate DEA-C01 troubleshooting practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.
Focused practice
Practice troubleshooting questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about troubleshooting
troubleshooting questions test whether you can apply the concept in context, not just recognise a definition.
How the topic appears in realistic exam-style scenarios.
Which detail in the question changes the correct answer.
How to eliminate plausible but wrong options.
How to connect the question back to the wider exam objective.
Watch out for
Common troubleshooting exam traps
- ▸Answering from memory before reading the full scenario.
- ▸Missing a constraint such as cost, availability, security, scope or command context.
- ▸Choosing a broad answer when the question asks for the most specific fix.
- ▸Ignoring why the wrong options are tempting.
Question index
All troubleshooting questions (1321)
Click any question to see the full explanation, or start a practice session above.
A data engineer needs to capture change data capture (CDC) events from an Amazon RDS for PostgreSQL database and stream them to Amazon S3 in near real-time. Which AWS service should be used?
Easy2A company stores application logs in Amazon S3 and wants to analyze them using Amazon Athena. The logs are in JSON format and are compressed with gzip. The data engineer needs to create an Athena table that can query these logs efficiently. The logs are stored in an S3 bucket with the prefix logs/year=2023/month=10/day=15/. The engineer wants to minimize query costs and improve performance. Which action should the engineer take?
Easy3A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3 in Parquet format. The replication is used for near-real-time analytics. Recently, the DMS task started failing with an error indicating insufficient memory. The source database is large (2 TB). What should a data engineer do to resolve this issue while minimizing changes to the existing architecture?
Hard4A data engineer is building an AWS Glue job that reads semi-structured JSON from Amazon S3 and must flatten nested arrays into relational columns before writing to Amazon Redshift. The transformation logic is complex and the engineer wants to unit test it locally without provisioning a cluster. Which Glue capability should the engineer use to develop and test this transformation logic?
Hard5A data engineer is optimizing an Amazon S3 data lake for cost and performance. The data lake contains large volumes of CSV files that are queried by Amazon Athena. The engineer wants to reduce query costs and improve query performance. Which TWO actions should the engineer take? (Choose two.)
Medium6A company stores sensitive financial data in an Amazon S3 bucket. The data engineering team must ensure that all data is encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys are rotated annually. The team also needs to audit key usage. Which solution meets these requirements?
Hard7Refer to the exhibit. This log snippet is from a failed AWS Glue job. The job processes a large dataset in memory. What is the MOST likely cause of the OutOfMemoryError?
Medium8A company has a large volume of CSV files in S3 that need to be transformed into Parquet using AWS Glue. The files are partitioned by date. The engineer wants to minimize costs by processing only new files each day. Which approach should be used?
Medium9A data engineer needs to ingest data from an on-premises Oracle database to Amazon S3 daily. The data volume is 500 GB per day, and the network bandwidth is 200 Mbps. The requirement is to minimize the impact on the source database and ensure data integrity. Which combination of AWS services should be used?
Medium10A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is JSON and each record is about 2 KB. The delivery stream is configured to buffer incoming data to 5 MB or 60 seconds, whichever comes first. The data engineering team notices that the S3 bucket contains many small files (average 2 MB), which makes subsequent processing inefficient. They need to reduce the number of small files without increasing the latency beyond 5 minutes. Which solution should they implement?
Medium11A data engineer is building an Amazon Redshift data warehouse. The cluster will ingest data from Amazon S3 using the COPY command. The engineer needs to ensure that the data is loaded in a way that maximizes query performance for future complex analytical queries. The data is currently stored as uncompressed CSV files in S3. Which action should the engineer take to optimize the load and subsequent query performance?
Medium12Which TWO features of Amazon S3 help protect data from accidental deletion or modification? (Choose two.)
Easy13A data engineer manages an Amazon S3 data lake that ingests millions of small JSON files daily from an IoT fleet. Query performance in Amazon Athena has degraded significantly, and each query scans far more data than expected. The engineer wants to reduce per-query cost and improve performance without changing the raw data. Which solution should the engineer implement?
Medium14A data engineer is monitoring an Amazon Redshift cluster and notices that the disk space usage is increasing rapidly. The engineer wants to reclaim space from deleted rows. Which command should the engineer run?
Easy15A data engineer is building an AWS Glue ETL job that reads data from Amazon S3 and must write the output partitioned by year, month, and day for efficient downstream querying in Amazon Athena. The engineer wants the job to create the partition folders and register them in the Glue Data Catalog automatically. (Choose two.)
Medium16A company is ingesting Apache logs from multiple web servers into AWS. The logs are sent via Amazon CloudWatch Logs to a subscription filter that delivers to a Lambda function. The Lambda function parses the logs and writes to Amazon S3. However, there is a significant backlog. Which THREE actions can reduce the backlog?
Hard17A data engineer is loading data from Amazon S3 into an Amazon Redshift cluster using the COPY command. The S3 bucket contains 500 Parquet files, each about 200 MB, in a single prefix. The COPY job is running slowly and consuming excessive cluster resources. The engineer wants to improve performance without changing the data format or the cluster size. Which action should the engineer take?
Medium18A data engineer is configuring an AWS Glue job that reads from an Amazon RDS for MySQL database and writes to Amazon S3. The security team requires that the data be encrypted in transit between AWS Glue and Amazon RDS. Which action should the engineer take to meet this requirement?
Medium19A data engineer manages an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. Both buckets are encrypted with SSE-KMS using customer managed keys. The Glue job execution role has permissions to read from the source bucket and write to the target bucket, and has kms:Decrypt permission on the source key, but the job fails with an error indicating it cannot write to the target bucket due to encryption. What is the MOST likely cause?
Medium20A data engineer notices that an S3 bucket policy allows access to a user from another AWS account, but the access is being denied. What could be the reason?
Hard21A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer wants to reduce query costs and improve performance for a table that is frequently queried with filters on a date column. The data is stored as uncompressed CSV files partitioned by year/month/day. Which action should the engineer take?
Easy22A company has a multi-account AWS environment with a centralized data lake in the Security account. Data producers in other accounts use AWS Glue to write data to S3 buckets in the Security account. The Security account uses AWS Lake Formation to manage permissions. The data engineer is setting up cross-account access so that users in the Producer account can query the data using Athena in their own account. The engineer has registered the S3 buckets and Data Catalog tables in Lake Formation. The IAM roles in the Producer account have the necessary permissions. However, when a user in the Producer account tries to query the table, they get an AccessDenied error. The error message indicates that the principal is not authorized to perform lakeformation:GetTable on the resource. What is the most likely cause?
Hard23A company wants to audit all data access events in their S3 buckets, including who accessed objects and from which IP address. Which AWS service should be used to capture these events?
Easy24A data engineering team stores clickstream events in an Amazon S3 bucket under the prefix s3://analytics/raw/. New objects arrive continuously, and the team wants Amazon Athena queries to scan only the events for the current day without scanning the entire prefix. The events are written as JSON files partitioned by year/month/day, but queries still scan all partitions because the partition metadata is not registered. Which action should the data engineer take to enable partition pruning in Athena?
Medium25A data engineer needs to transform JSON data from Amazon S3 into Parquet format using AWS Glue. The data contains nested fields. Which Glue feature should the engineer use to define the schema and handle the nested structure?
Easy26A data engineer needs to run a transformation on a small dataset of 500 MB stored in Amazon S3 and load the result into Amazon Redshift. The transformation logic is simple column renaming and filtering. The engineer wants to minimize operational overhead and avoid managing servers. Which approach is most appropriate?
Easy27Which TWO of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose TWO.)
Medium28A financial services company stores trade records in Amazon DynamoDB. An application performs many reads per second for individual trades by tradeId and also runs a nightly analytics job that must read every trade for a given trading day. The table uses tradeId as the partition key only. The nightly job currently performs a full table scan and is slow and expensive. The company wants to optimize the nightly access pattern with minimal application changes. Which change should the data engineer make?
Hard29A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The engineer notices that some data records are missing from S3, and the Firehose delivery stream metrics show increased DeliveryToS3.DataFreshness. The engineer needs to ensure all records are delivered. Which action should the engineer take?
Medium30A company needs to audit all API calls made in their AWS account, including actions performed by the root user. Which AWS service should be used?
Easy31A data engineer needs to ingest data from an on-premises Apache Kafka cluster into Amazon S3. The data volume is about 10 TB per day. The engineer wants to set up a managed Kafka connector. Which AWS service should they use?
Medium32A company is building a data lake on S3 and needs to ingest data from on-premises Oracle database. The data is 5 TB and changes incrementally. The ingestion must capture changes in near real-time (less than 1 minute latency) and be cost-effective. Which approach should be used?
Hard33A data engineer notices that an Amazon Redshift cluster is experiencing slow query performance. The engineer suspects that tables are not properly sorted. Which diagnostic query should the engineer run to identify unsorted rows?
Medium34A company uses Amazon DynamoDB for a gaming application. They need to store player session data that expires after 24 hours. Which DynamoDB feature should they use to automatically delete expired items?
Easy35A data engineer needs to ingest data from a SaaS application that sends webhooks in JSON format. The data must be stored in S3 for batch analysis. Which AWS services can receive the webhooks and store the data in S3 with minimal custom code? (Choose TWO.)
Easy36A company is using Amazon Redshift for data warehousing. The data engineer notices that the STL_ALERT_EVENT_LOG table shows many 'missing statistics' alerts. What is the best course of action to address this issue?
Hard37A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 data source, applies a filter transformation, and writes to Amazon S3 in Parquet. The engineer notices that the job is reading all files in the prefix, including files that do not match the expected schema, causing job failures. Which action should the engineer take to ensure only valid files are processed?
Medium38A data engineer is designing a data transformation pipeline using AWS Glue. The source data is in Amazon S3 in Parquet format, and the transformed output must be written to another S3 bucket in Parquet format partitioned by year, month, day. The pipeline should handle incremental updates efficiently. Which three features should the engineer use? (Choose THREE.)
Hard39A company is using Amazon RDS for MySQL and needs to reduce read latency for a global user base. Which AWS feature should be implemented?
Medium40A company uses AWS Glue crawlers to populate the Data Catalog from data in Amazon S3. The crawler fails to update the schema when new columns are added to the CSV files. What is the most likely cause?
Medium41A company runs a multi-AZ Amazon RDS for PostgreSQL instance. They need to run a one-time analytical query that will take several hours and consume significant I/O. The query should not impact the primary workload. What should the data engineer do?
Medium42A company stores regulated records in an Amazon S3 bucket and must prove that individual objects cannot be deleted or overwritten for 365 days after creation, even by the account root user. The compliance team also needs to retain the ability to delete the bucket itself after the retention window expires. Which configuration meets these requirements?
Hard43A data engineer is building an AWS Lake Formation governed data lake. The security team wants to grant a group of analysts access to only the non-sensitive columns of a table in the Data Catalog, while denying access to columns containing Social Security numbers. The analysts use Amazon Athena to query the data. Which Lake Formation permission model should the data engineer use?
Medium44A data engineer is designing a multi-region disaster recovery plan for an Amazon DynamoDB table. The table stores critical user profile data and must have a Recovery Point Objective (RPO) of less than 1 minute and a Recovery Time Objective (RTO) of less than 5 minutes. Which solution meets these requirements?
Hard45A data engineer needs to transform JSON data from an S3 bucket using AWS Glue. The JSON contains nested arrays and objects. Which Glue transform is best suited for flattening nested structures?
Easy46A data engineer is troubleshooting an issue where an AWS Glue ETL job fails when trying to read data from an S3 bucket encrypted with SSE-KMS. The job has an IAM role that includes `kms:Decrypt` permission. What is the most likely reason for the failure?
Medium47A company has an Amazon RDS for PostgreSQL DB instance with a large table that is frequently updated. The data engineer needs to reduce storage costs by archiving old records that are no longer accessed. The archived records must be retained for 7 years due to compliance requirements. Which solution is MOST cost-effective?
Medium48Refer to the exhibit. A data engineer runs this CLI command to check an object's metadata. The engineer wants to verify if the object is eligible for lifecycle transition to S3 Glacier based on its age. What additional information is needed?
Easy49A data engineer is optimizing an AWS Glue ETL job that reads a large dataset from Amazon S3 and writes to Amazon Redshift. The job currently runs slowly and consumes many DPUs. The engineer wants to improve performance and reduce cost. Which two actions should the engineer take? (Choose two.)
Hard50A data engineering team is building a data lake on Amazon S3. They need to catalog data and make it queryable by Amazon Athena and Amazon Redshift Spectrum. The data arrives in multiple formats and the schema evolves frequently. Which TWO actions should the team take to support schema evolution and efficient querying? (Choose two.)
Hard51A data pipeline uses AWS Glue to read from an Amazon S3 bucket containing millions of small CSV files (each < 1 MB). The ETL job is slow. Which optimization would most improve performance?
Hard52A data engineer is implementing a CDC (Change Data Capture) pipeline from a relational database to Amazon S3 using AWS Database Migration Service (DMS). Which TWO configurations are required for continuous replication?
Hard53A data engineer needs to ensure that an Amazon Redshift cluster encrypts all data at rest. Which setting must be enabled when creating the cluster?
Easy54A data engineer is troubleshooting an Amazon Redshift cluster that is not allowing connections from a specific IP range. The engineer verified that the cluster's security group allows inbound traffic from the IP range. What is the next step to resolve the issue?
Hard55A company uses Amazon Redshift to store customer data. The security team requires that all queries are logged for auditing purposes. Which step should be taken to meet this requirement? (Select ONE.)
Medium56A company uses Amazon DynamoDB as the primary data store for a gaming application. The application experiences sudden spikes in traffic. The data engineer notices that write requests are throttled during peak times. The partition keys are well-distributed. What should the data engineer do to reduce throttling?
Hard57A data engineer stores sensitive records in an Amazon S3 bucket. The security team wants to guarantee that every object is encrypted before it is written to disk and that the bucket automatically rejects any unencrypted PUT request, regardless of which IAM principal sends it. Which configuration should the data engineer apply?
Easy58A company uses AWS Lake Formation to manage data lake permissions. A data analyst cannot query a table in Athena, although the table appears in the catalog. The analyst has IAM permissions to run Athena. What is the MOST likely cause?
Medium59A company runs a data pipeline that uses AWS Glue to process data from an Amazon DynamoDB table and write results to Amazon S3. The Glue job runs on a schedule every hour. Recently, the job started failing intermittently with 'ProvisionedThroughputExceededException' errors from DynamoDB. What is the BEST solution?
Medium60Refer to the exhibit. A data engineer is troubleshooting an AWS Lambda function that processes data from Amazon S3. The function is triggered by S3 events, but no logs appear in CloudWatch Logs. The engineer runs the AWS CLI command shown. What is the MOST likely reason for the missing logs?
Medium61A data engineer is designing a data lake on Amazon S3 that will store sensitive financial data. The security team requires that access to the data be audited, that data be encrypted at rest with customer-managed keys, and that the engineer be able to identify which IAM principals accessed specific objects. Which TWO AWS services or features should the engineer use to meet these requirements? (Choose two.)
Hard62A data engineer manages an AWS Glue ETL job that processes JSON files from Amazon S3 and writes to Amazon Redshift. The job fails with the error 'Unable to find a suitable JDBC driver in the classpath'. The engineer has included the Redshift JDBC driver as a job parameter in the Glue job configuration. Which step should the engineer take to resolve the error?
Medium63A data engineer needs to load data from an Amazon S3 bucket into an Amazon Redshift cluster as part of an ETL pipeline. The source files are already in Parquet format and the engineer wants the fastest load with minimal transformation. Which Redshift load method should the engineer use?
Easy64A data pipeline uses AWS Glue to process data from Amazon S3. The job fails with an 'OutOfMemoryError' during the transformation phase. Which action should the data engineer take to resolve this issue?
Medium65A data engineer is setting up an AWS Glue job to process data from an Amazon S3 bucket. The job fails with an 'Access Denied' error. Which TWO IAM permissions are MOST likely missing from the Glue job's IAM role?
Easy66A data engineer needs to transform data in Amazon S3 using SQL statements without managing any infrastructure. The transformations are simple projections and filters, and the engineer wants the results written back to S3 in Parquet. Which AWS service should be used?
Medium67A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The migration must have minimal downtime. Which service should be used to perform the migration?
Medium68A company stores raw customer records in an Amazon S3 bucket and processes them with AWS Glue. A governance requirement states that a specific tag named DataClass must exist on every catalog table, and any table missing that tag must not be queryable. Where should the data engineer enforce this requirement with the least operational effort?
Easy69A company needs to transform JSON data from an Amazon S3 bucket into Parquet format and load it into an Amazon Redshift cluster. The transformation includes joining with a reference table stored in Amazon RDS. Which AWS service is BEST suited for this task?
Medium70A company uses Amazon Kinesis Data Firehose to ingest application logs into an Amazon S3 bucket. The logs are in JSON format. The data engineering team wants to convert the logs from JSON to Parquet format before landing in S3. What is the most cost-effective way to achieve this?
Medium71A company is storing large amounts of log data in Amazon S3. The data is accessed frequently for the first 30 days, then rarely after that. The company wants to automatically transition the data to a lower-cost storage class after 30 days. Which S3 feature should the data engineer use?
Easy72A data engineer is designing a data lake on S3 and needs to ensure that data is encrypted at rest using customer-managed KMS keys. The engineer also needs to audit all access to the KMS keys. Which combination of services should be used?
Medium73Refer to the exhibit. A data engineer is using a Kinesis Data Stream with 2 shards. The producer uses a partition key that is the user ID (a UUID). The consumer is falling behind. Which change would improve throughput?
Hard74A company needs to ingest data from a MySQL database into Amazon S3 using AWS DMS. The data changes frequently and the requirement is to capture changes in near real-time. Which THREE configurations are necessary?
Hard75A company uses Amazon Kinesis Data Streams with a Lambda consumer. The Lambda function is failing with 'ProvisionedThroughputExceededException' when writing to a DynamoDB table. Which action should the data engineer take to resolve this without losing data?
Hard76A data pipeline uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The delivery stream is configured with a buffer interval of 60 seconds and a buffer size of 5 MB. The data arrives at an average rate of 2 MB per second. What is the expected time interval between S3 writes?
Hard77A healthcare company is building a data pipeline to ingest electronic health records (EHR) from hospitals. The data is sent as JSON files via SFTP to an on-premises server. The company wants to move this data to AWS using AWS Transfer Family (SFTP) and then process it with AWS Glue. Data sovereignty regulations require that all data remain within the EU (Frankfurt) region. The pipeline must detect when a new file arrives and start the Glue job automatically. The engineer has set up an AWS Transfer Family server in Frankfurt, and files are uploaded to an S3 bucket in the same region. However, the Glue job is not triggering automatically. The engineer needs to implement automated triggering. What should the engineer do?
Hard78A data engineer needs to enforce that all data in an Amazon S3 bucket is encrypted at rest. Which of the following can be used to achieve this? (Choose TWO.)
Easy79A data engineer needs to store large volumes of infrequently accessed compliance data in Amazon S3 for 10 years. The data must be retrievable within 12 hours if required for audits. The engineer wants the most cost-effective storage solution. Which S3 storage class should be used?
Easy80A company ingests streaming data into Amazon Kinesis Data Streams. Producers write records using the PutRecords API with explicit partition keys based on customer ID. A data engineer observes that a few shards are consistently at 100 percent write throughput while others are underutilized, causing throttling. Which action should the engineer take to distribute the load more evenly?
Medium81A data engineer is using Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to orchestrate a data pipeline. The engineer needs to ensure that the Airflow environment can access an Amazon S3 bucket to read and write data. The S3 bucket is in the same AWS account and Region as the MWAA environment. Which configuration is required to allow MWAA to access the S3 bucket?
Medium82A data engineer is responsible for an Amazon Redshift cluster that ingests data continuously from Amazon Kinesis Data Streams. The engineer needs to ensure that the raw streaming data is immediately queryable in Redshift with minimal latency. Which approach should the engineer take?
Medium83A company needs to ingest data from multiple on-premises databases into Amazon S3 for analytics. The databases include Oracle, MySQL, and PostgreSQL. The data must be continuously replicated with minimal latency. Which AWS service should be used?
Easy84A data engineer is configuring an Amazon S3 lifecycle policy to transition objects to S3 Glacier Deep Archive after 90 days. The bucket receives new objects daily. The engineer wants to ensure that objects are not deleted before 90 days. Which lifecycle action should be used?
Easy85A company wants to securely store database credentials used by a Lambda function. Which AWS service should be used to store and rotate the credentials automatically?
Easy86A data engineer has an AWS Glue ETL job that processes JSON files from Amazon S3. The job currently uses the DynamicFrame method to write output to Amazon Redshift. The engineer needs to improve write performance by using a staging Amazon S3 bucket and parallel COPY operations. Which AWS Glue connection option should the engineer configure?
Medium87A company uses AWS Glue to run ETL jobs on data stored in S3. The data is encrypted with SSE-KMS. The Glue job fails with an 'AccessDenied' error when trying to read the data. What is the MOST likely cause?
Medium88Refer to the exhibit. A data engineer runs a Glue job manually and receives a ThrottlingException. The engineer checks the job run history and sees a previous failure with the same error. What is the MOST likely cause of the throttling, and which solution is MOST appropriate?
Hard89A data engineer maintains an AWS Glue ETL job that processes millions of small JSON files stored in Amazon S3. The job's runtime has increased significantly, and CloudWatch logs show many small executor tasks and frequent garbage collection. The engineer wants to improve job performance by reducing the number of small files processed per task. Which action should the engineer take?
Medium90A data engineer needs to schedule an AWS Glue extract, transform, and load job to run every day at 02:00 UTC and trigger a dependent Amazon Redshift stored procedure only after the Glue job succeeds. The engineer wants a managed orchestration option that avoids provisioning servers. Which approach should the engineer use?
Easy91A data engineer is troubleshooting a slow-running Amazon Athena query on a large dataset stored in S3. The query scans many small files. Which TWO actions can improve query performance?
Medium92A company needs to store application log files for 90 days for compliance. The logs are generated continuously and are rarely accessed after 30 days. The data engineer must minimize storage costs. Which storage solution should the engineer choose?
Easy93A company is using Amazon Kinesis Data Streams with a Lambda consumer to process clickstream data. The data rate is high and the Lambda function is falling behind, resulting in increased processing latency. What is the MOST effective way to improve throughput?
Medium94A company wants to enforce encryption in transit for data moving between an EC2 instance and an S3 bucket. Which TWO methods can achieve this? (Choose 2)
Easy95Which TWO actions are recommended for securing data at rest in Amazon S3? (Choose two.)
Medium96A data engineer needs to set up a data catalog for a new data lake in AWS Glue. The data resides in S3 in Parquet format. The engineer wants to ensure that the schema is automatically detected and updated when new columns are added to the data. Which configuration should the engineer use?
Medium97A data engineer manages an AWS Glue ETL job that processes millions of small JSON files in Amazon S3. The job is slow and often fails with an OutOfMemory error on the driver. The engineer wants to improve performance without changing the output format. Which solution should the engineer implement?
Medium98A data team runs a daily AWS Glue ETL job that processes data from an Amazon Redshift cluster and writes results to Amazon S3. The job completes successfully but takes 2 hours longer than expected. The job uses the JDBC connection to Redshift. The Redshift cluster is 4 dc2.large nodes. The Glue job has 10 workers of type G.1X. Which change would MOST likely reduce the job duration?
Hard99A data engineer is responsible for monitoring an AWS Glue ETL job that runs daily. The job reads data from an Amazon S3 bucket and writes to an Amazon Redshift table. The engineer wants to receive an alert if the job fails or if it takes longer than expected to complete. Which AWS service should the engineer use to set up these alerts with the LEAST operational overhead?
Easy100A company is using AWS Glue to catalog data in Amazon S3. The data is stored in CSV format, but the schema is not consistent across all files. Which TWO actions can the company take to handle schema evolution and ensure the Glue Data Catalog is up to date? (Choose TWO.)
Easy101A company uses AWS DMS to continuously replicate data from an on-premises SQL Server to Amazon Aurora MySQL. The replication lag is increasing. Which THREE actions can reduce the lag? (Choose three.)
Hard102A data analyst needs to query a large Amazon S3 bucket containing CSV files using Amazon Athena. The bucket has millions of small files (less than 1 MB each). The analyst reports that queries are very slow and often time out. The data is partitioned by date and the partition columns are defined in the table. What is the most effective way to improve query performance?
Easy103A company runs a data warehouse on Amazon Redshift. Queries are slow, and the team suspects data distribution is skewed. Which approach would best help identify distribution skew?
Medium104An organization is using AWS Glue to process sensitive data. The data is stored in S3 with server-side encryption using AWS KMS (SSE-KMS). The Glue job fails with an error indicating that it cannot read the data. The IAM role used by Glue has the following policy. What is missing?
Hard105A company uses Amazon Kinesis Data Streams to ingest real-time data. The compliance team requires that all data in the stream be encrypted at rest. Which configuration should be enabled?
Medium106A data engineer needs to ingest JSON data from an on-premises relational database into Amazon S3 every hour. Which AWS service should be used to set up a scheduled, incremental data transfer?
Easy107A data engineer is troubleshooting an Amazon DynamoDB table that has frequent throttling exceptions for write requests. The table has auto scaling enabled. What is the most likely cause?
Hard108Match each AWS data analytics service to its primary function.
Medium109A company is running a data warehouse on Amazon Redshift. The data engineering team notices that query performance has degraded over time. They suspect that data distribution is causing excessive data movement between nodes. The table is joined frequently on the customer_id column. Which column should be chosen as the distribution key to optimize join performance?
Medium110A company uses Kinesis Data Firehose to deliver streaming data to S3. They need to transform the data by adding a timestamp and removing sensitive fields. Which TWO approaches can achieve this?
Easy111A data engineer created the IAM policy shown in the exhibit. The engineer then attempts to upload an object to 'my-bucket' using the AWS CLI with the command: aws s3 cp file.txt s3://my-bucket/ --sse aws:kms. The upload fails with an 'AccessDenied' error. What is the most likely cause?
Hard112A data engineer needs to audit all access to an Amazon S3 bucket containing sensitive data. The audit must capture who accessed the bucket, from which IP address, and what actions were performed. Which AWS service should be enabled?
Medium113A data engineer needs to ensure that data stored in Amazon S3 is protected against accidental deletion and that the data cannot be altered for a period of 7 years to meet compliance requirements. The engineer must also ensure that the data remains encrypted at rest using SSE-KMS. Which two actions should the engineer take? (Choose two.)
Medium114A data engineer needs to share a dataset from an S3 bucket in Account A with another AWS account (Account B). The data must remain encrypted at rest with KMS. Which steps are required?
Easy115A company wants to ingest real-time clickstream data from a website into Amazon S3 with minimal code. The data should be delivered within 60 seconds of generation. Which AWS service should be used?
Easy116A data engineer is setting up a data pipeline using AWS Glue. The engineer wants to monitor job failures and receive notifications. Which TWO services can be used together for this purpose?
Easy117A company is using Amazon Redshift for analytics and needs to ensure that all data is encrypted at rest. The current cluster does not have encryption enabled. What is the most efficient way to enable encryption?
Medium118A data engineer is designing a data lake on AWS using Amazon S3. The data consists of CSV files generated by IoT devices. The data is accessed by multiple analytics jobs, and the engineer needs to ensure that new files are immediately visible to all consumers after writing. What S3 consistency model applies?
Easy119A data engineer reviewed the S3 lifecycle policy shown in the exhibit. The engineer notices that objects under the 'logs/' prefix are being deleted after 365 days. The business requirement is to retain logs for at least 5 years. What should the engineer change in the lifecycle policy?
Medium120A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The Flink application reads from a Kinesis Data Streams source, performs aggregations, and writes results to Amazon S3. The application is experiencing high checkpoint failures, and the processing lag is increasing. The data volume is 50 MB/s with an average record size of 1 KB. Which TWO actions would improve checkpoint reliability and reduce lag? (Choose TWO.)
Hard121A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job must be able to handle a large number of small files efficiently and minimize the number of output files to improve downstream query performance. Which two actions should the engineer take? (Choose two.)
Medium122A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a nested folder structure. The engineer wants to automatically discover the schema and update the AWS Glue Data Catalog as new files are added. Which solution should the engineer use?
Medium123A data engineer is monitoring an Amazon EMR cluster and notices that the cluster is running out of disk space on the core nodes. Which action can be taken to resolve this issue?
Easy124A data engineer must ensure that all objects written to an S3 bucket by an AWS Glue ETL job are encrypted with a customer-managed AWS KMS key, even if the job does not explicitly specify encryption parameters. The bucket policy already denies unencrypted PUT requests. Which configuration will enforce the required encryption with the LEAST operational overhead?
Medium125A data engineer is using AWS Step Functions to orchestrate a daily ETL workflow that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The workflow occasionally fails and the engineer needs to troubleshoot and recover. Which TWO actions should the engineer take to identify the failure and resume from the failed step? (Choose two.)
Hard126A financial services company processes real-time stock trade data. They use Amazon Kinesis Data Streams with a shard count of 5, each shard receiving about 500 records per second. The consumer application uses the Kinesis Client Library (KCL) with DynamoDB for checkpointing. Lately, some records are being processed multiple times. What is the most likely cause?
Hard127A data engineer is using AWS Step Functions to orchestrate a daily ETL pipeline that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The pipeline occasionally fails with the error 'States.TaskFailed' from the Glue job, but the Glue job's own logs show that it completed successfully. The Step Functions execution history shows that the Glue job task timed out after 15 minutes, while the Glue job actually ran for 18 minutes. The Step Functions state machine uses the optimized Glue service integration with a TaskTimeout of 900 seconds. Which change will allow the pipeline to complete successfully without reducing the Glue job's runtime?
Hard128A company needs to ingest streaming data from thousands of IoT devices. The data must be processed in real-time and stored in Amazon S3. Which TWO services should be used together?
Medium129A data engineer needs to securely store database credentials used by a Lambda function. The solution must automatically rotate the credentials every 90 days. Which AWS service should the engineer use?
Easy130A company is building a data pipeline that ingests streaming data from IoT devices. The data must be stored in a durable, scalable, and cost-effective manner for batch processing. Which TWO AWS services should be used together?
Easy131A data engineer is running an AWS Glue ETL job that reads from an Amazon RDS MySQL database and writes to Amazon S3. The job fails with a 'Communications link failure' error. The security group for the RDS instance allows inbound traffic from the Glue job's security group. What is the most likely cause of the failure?
Easy132A company is ingesting data from multiple sources into S3 using AWS Glue. The data engineer notices that the Glue job is failing with an OutOfMemory error. Which step should be taken to resolve this issue?
Medium133A data engineering team is using AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration must have minimal downtime and needs to capture ongoing changes after the full load. Which THREE resources are required for this task? (Choose three.)
Medium134A data engineer is using AWS Glue to process data from an Amazon Kinesis Data Stream. The Glue job is configured to run every 15 minutes and uses job bookmarks to track processed data. Recently, the job started reprocessing old data, leading to duplicate records in the target Amazon S3 bucket. The engineer verifies that the job bookmark is enabled and the job is not being run manually. Which TWO actions should the engineer take to resolve the duplicate processing issue? (Choose two.)
Hard135A company uses AWS Glue ETL to process data from Amazon S3 and write results to Amazon Redshift. The job fails with a memory error when processing large files. Which action should the data engineer take to resolve this issue?
Medium136A company uses AWS Glue to run ETL jobs on a schedule. Recently, a job failed with the error: 'AnalysisException: cannot resolve '`column_name`' given input columns: ...'. The job reads from an Amazon S3 source that has a schema defined in the AWS Glue Data Catalog. What is the MOST likely cause?
Medium137A company is using Amazon S3 to store large amounts of archival data. The data is accessed infrequently but must be immediately retrievable when needed. Which storage class is the most cost-effective choice?
Medium138A data engineer is setting up an Amazon Kinesis Data Analytics application to process streaming data from a Kinesis data stream named "input-stream". The application uses a reference data source from an S3 bucket. The engineer has attached the IAM policy shown in the exhibit to the application's IAM role. When starting the application, the engineer receives an 'AccessDeniedException' error. Which additional permission is required?
Hard139A data engineer needs to run an AWS Glue for Apache Spark ETL job that joins a 40 GB Amazon S3 Parquet dataset with a small 8 MB reference lookup table stored as CSV in Amazon S3. The reference table is read on every join and the job's executors are spending a large amount of shuffle time on the join. The reference table changes only once per month. Which approach MOST efficiently reduces shuffle overhead in the Glue job?
Medium140A company is designing a data lake on Amazon S3. Which TWO strategies improve query performance for Amazon Athena?
Medium141A company uses AWS KMS to encrypt sensitive data stored in S3. To meet compliance requirements, they need to ensure that the encryption keys are automatically rotated every year. Which type of KMS key should they use?
Hard142A company uses AWS Glue to run ETL jobs daily. The jobs consume data from an Amazon RDS for MySQL database and write results to Amazon S3. The company wants to minimize the impact on the source database during extraction. Which THREE actions should the data engineer take to achieve this? (Choose THREE.)
Medium143A data engineer is troubleshooting an AWS Glue ETL job that fails with the error 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large number of small files in Amazon S3. Which action would MOST effectively resolve the issue?
Hard144A data engineer is troubleshooting a failed AWS Glue job that reads from an Amazon RDS for MySQL table. The error message indicates 'java.sql.SQLException: No suitable driver'. What is the most likely cause?
Medium145A company is using Amazon Redshift for its data warehouse. The data engineering team needs to improve query performance for a large fact table that is frequently joined with multiple dimension tables. Which THREE strategies should be considered?
Hard146A data engineer needs to store sensitive data in Amazon S3 and automatically classify the data using a managed service. The data is uploaded via an S3 bucket. Which AWS service can automatically detect and classify sensitive data?
Medium147A company stores sensitive data in Amazon S3. The security team requires encryption at rest and that the encryption keys are managed by the company using AWS KMS. The data is frequently accessed by multiple AWS services. Which THREE steps should be taken to meet these requirements?
Hard148An AWS Glue job that performs data transformation on large Parquet files in Amazon S3 is taking a long time to complete. The job uses the default number of DPUs. Which change would most likely improve the job's performance?
Medium149Which THREE AWS services can be used to centrally manage and govern data across multiple AWS accounts? (Select THREE.)
Easy150A company is using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data must be transformed from JSON to Parquet format before landing in S3. The transformation logic is simple: convert the JSON schema to Parquet. Which approach meets the requirements with the least operational overhead?
Medium151Refer to the exhibit. A Lambda function named 'IngestionProcessor' is failing. The engineer checks CloudWatch Logs and sees the log group exists but storedBytes is 0. Why might the logs show no data?
Easy152A data engineer is designing a data pipeline that processes PII data using AWS Glue and stores results in S3. Which TWO actions should be taken to protect the data? (Choose 2)
Medium153A data engineer is designing a data pipeline that ingests data from an on-premises system into Amazon S3 using AWS Transfer Family. The data must be encrypted at rest using a customer-managed key in AWS KMS. The S3 bucket policy must allow only encrypted connections. Which policy condition should be used?
Hard154A company is using Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data is delivered in 5-minute intervals. The company wants to reduce the delivery frequency to 1 minute to get data faster. Which parameter should be changed in the Firehose delivery stream configuration?
Easy155A company uses AWS Lake Formation to manage access to a data lake in Amazon S3. A data engineer needs to grant a specific IAM role access to only the columns containing non-sensitive data in a table, while hiding columns with personally identifiable information (PII). The engineer has already registered the S3 bucket and the table in Lake Formation. What should the engineer do to meet this requirement?
Hard156A company uses Amazon RDS for PostgreSQL. The data engineer needs to ensure that the database is automatically backed up and that backups are retained for 35 days. What is the simplest way to achieve this?
Easy157A data engineer is designing a streaming pipeline using Amazon Kinesis Data Streams with a shard count of 10. The incoming data rate is 1 MB/second. The consuming application uses the Kinesis Client Library (KCL) with a single worker. What is the most likely performance bottleneck?
Hard158A company wants to ingest real-time clickstream data from a website into Amazon S3 with a maximum latency of 60 seconds. The data volume peaks at 500 MB/s. Which service should they use to buffer and deliver the data to S3?
Easy159A data engineer needs to transform data in an AWS Glue job using a custom Python library that is not available by default. The library is packaged as a .whl file stored in Amazon S3. The engineer wants the Glue job to use this library without modifying the job script to install it at runtime. What should the engineer do?
Medium160A company is using Kinesis Data Firehose to deliver data to an S3 bucket. The delivery stream is failing with 'S3 bucket access denied' errors. The bucket policy allows the Firehose service principal. What could be the issue?
Medium161A data engineer is designing an Amazon Redshift data warehouse for a high-traffic analytics workload. The engineer needs to ensure fast query performance and minimize data movement. Which THREE design decisions should be made? (Choose THREE.)
Hard162A data engineer is monitoring an AWS Glue ETL job that processes data from an S3 bucket and writes to a Redshift table. The job completes successfully but takes longer than expected. The engineer notices that the job uses 10 DPUs and the data size is 500 GB. The job runs in standard mode. Which change would MOST reduce job duration?
Medium163A company uses Amazon Kinesis Data Firehose to ingest log data from web servers into Amazon S3. The data is in JSON format and each record is approximately 2 KB. The delivery stream is configured to buffer incoming records for 60 seconds or 5 MB, whichever comes first. The company notices that the data in S3 is delayed by up to 5 minutes during peak hours. Which action would most effectively reduce the delivery latency?
Hard164A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow due to many small inserts. Which technique would improve write performance?
Medium165A company uses AWS Glue to process data from multiple S3 buckets. The Glue job runs daily and reads data from a bucket that contains millions of small files (each < 1 MB). The job has been running for hours and is often close to the 8-hour timeout limit. Which optimization would MOST reduce the job's runtime?
Hard166A data engineer is configuring an AWS Glue ETL job that processes data from an Amazon S3 bucket. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS. The engineer needs to ensure that the Glue job uses HTTPS endpoints when reading from and writing to S3. Which action should the engineer take?
Easy167A company uses Amazon DynamoDB as its primary data store for a web application. The application experiences high latency during peak hours. The data engineer notices that the table has a large number of items with the same partition key. Which DynamoDB feature should the engineer use to improve performance?
Easy168A data engineer is designing a data pipeline that ingests data from multiple sources into Amazon S3, then processes it with AWS Glue and loads it into Amazon Redshift. Which THREE practices should be implemented to ensure data quality?
Hard169A company needs to store JSON documents that are accessed by a key-value pattern. The data is 500 GB and requires single-digit millisecond latency. Which AWS database is most suitable?
Easy170A data engineer is ingesting records from Amazon Kinesis Data Streams into Amazon S3 using AWS Lambda as the consumer. Each stream shard delivers up to 1,000 records per second, and the Lambda function writes each record as an individual small object, causing many tiny S3 files and high PUT costs. The engineer wants fewer, larger objects while keeping near-real-time delivery. What should the engineer do?
Medium171A company needs to monitor and record all changes to IAM policies in their AWS account. Which AWS service should be used?
Medium172A company uses Amazon Kinesis Data Firehose to ingest data into an S3 bucket. The data is in JSON format and the team wants to convert it to Parquet before storage. Which TWO configurations are required?
Medium173A company uses AWS Glue DataBrew for data preparation. The data source is an S3 bucket with millions of small CSV files (each < 1 MB). The DataBrew project takes a long time to load the sample data. What is the most likely cause and solution?
Hard174A company uses Amazon Redshift for its data warehouse and needs to enforce column-level security on sensitive columns. Which TWO approaches can achieve this?
Hard175Which TWO methods can be used to enforce least-privilege access to an Amazon S3 bucket? (Choose two.)
Easy176A company wants to centrally manage encryption keys for multiple AWS services and automatically rotate them every year. Which AWS service should be used?
Medium177A team uses Amazon Redshift for analytics. They notice that some queries are slow and the system shows high disk usage. The team wants to improve query performance without adding more nodes. Which action should they take first?
Medium178Which TWO actions can help improve query performance in Amazon Redshift? (Choose two.)
Medium179A data engineering team is designing a data lake on Amazon S3. They need to store raw data in its original format and transformed data in Parquet. The data is accessed by multiple analytics services, including Amazon Athena and Amazon Redshift Spectrum. Compliance requirements mandate that all data be encrypted at rest with AWS KMS and that the encryption keys be rotated every 90 days. Which S3 bucket configuration meets these requirements?
Hard180A data engineer needs to ensure that all data stored in an Amazon S3 bucket is encrypted at rest using a customer managed key in AWS KMS. The engineer also needs to enforce that any attempt to upload an object without specifying the correct KMS key is denied. Which combination of actions should the engineer take?
Medium181A company is using AWS KMS to encrypt data in Amazon S3. The security team wants to ensure that only specific IAM roles can decrypt the data. Which TWO steps should the data engineer take? (Choose two.)
Hard182A data engineer needs to run a transformation that processes semi-structured JSON records already stored in Amazon S3 and write the results back to S3 in Parquet format. The team prefers a serverless, Apache Spark-based approach with minimal infrastructure management and wants to use the AWS Glue Data Catalog for metadata. Which approach should the engineer use?
Easy183A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is running but the engineer observes that some tables are not being replicated. The DMS task logs show no errors, and the task status is 'Running'. Which action should the engineer take to identify the missing tables?
Easy184A company wants to monitor and alert on unauthorized API calls in their AWS account. Which AWS service should be used to detect and notify on such events?
Medium185A data engineer runs the describe-stream command and sees the output above. The stream has a retention period of 24 hours. The engineer needs to ensure that consumers can replay data for up to 7 days. Which action is required?
Hard186A data engineer needs to ingest streaming data from a social media API into Amazon S3 for batch analytics. The data arrives at a rate of 500 records per second. Which service should be used to capture the stream?
Easy187A company is streaming IoT data from thousands of devices into Amazon Kinesis Data Streams. The data must be transformed in real time before being stored in Amazon S3. Which service should be used to perform the transformation as the data streams through Kinesis?
Medium188Which TWO AWS services can be used as sources for AWS Glue ETL jobs? (Choose two.)
Easy189A company is using Amazon Redshift for data warehousing. They need to ensure that data is encrypted at rest and in transit. Which TWO configurations are required to meet these requirements?
Medium190A data engineer is troubleshooting an AWS Glue ETL job that fails with an access denied error when writing to an S3 bucket. The Glue job uses an IAM role that has an S3 bucket policy attached. The bucket policy denies access to any principal that does not use server-side encryption. What is the most likely cause of the failure?
Hard191A data engineer is designing a data lake on Amazon S3. The team wants to optimize query performance and reduce storage costs for a large dataset of JSON logs that are queried frequently by Amazon Athena. The logs are currently stored as uncompressed JSON files, each around 1 GB, in a single prefix. The engineer needs to improve query performance and reduce costs without changing the data format. Which action should the engineer take?
Medium192A company is using AWS Glue to catalog data stored in Amazon S3. The data is partitioned by year, month, and day. A data analyst reports that new partitions are not automatically discovered by the Glue crawler. The crawler runs on a schedule every hour. What is the MOST likely reason for the missing partitions?
Easy193Match each AWS service to its primary purpose in data engineering.
Medium194A data engineer is using AWS Step Functions to orchestrate a daily pipeline that runs several AWS Glue jobs in sequence. One Glue job intermittently fails due to a transient Amazon S3 503 error. The engineer wants the state machine to automatically retry only that Glue job up to three times with exponential backoff, without retrying the other jobs. What should the engineer do?
Easy195A company runs a real-time analytics platform on Amazon ECS that ingests streaming data from Amazon Kinesis Data Streams, processes it, and stores results in Amazon DynamoDB. The data volume spikes unpredictably, causing DynamoDB to throttle write requests. The application uses on-demand capacity mode. The data engineer notices that the throttling occurs on a specific partition due to a hot key. The hot key is a customer ID that receives a disproportionate number of writes. The application cannot change the partition key design immediately. The engineer needs to reduce throttling while maintaining low latency. Which solution is most effective?
Hard196A company is using an Amazon RDS for PostgreSQL database to store personally identifiable information (PII). The security team wants to ensure that database administrators cannot view the plaintext PII data. Which solution should a data engineer implement?
Hard197A data engineer needs to run a one-time transformation on a 500 GB dataset stored in Amazon S3. The transformation is written in Python and uses pandas, which cannot handle the full dataset in memory. The engineer wants a serverless option that can parallelize the work without managing servers. Which AWS service should the engineer use?
Easy198A company uses Amazon DynamoDB as the primary data store for a web application. The application experiences occasional throttling on write requests. The data engineer needs to implement a solution that handles throttling gracefully without losing data. Which approach should the engineer use?
Easy199A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Redshift. The migration is successful, but after a few days, data in Redshift becomes inconsistent with the source due to ongoing changes. The company needs to keep Redshift synchronized with minimal latency. Which approach should the data engineer use?
Hard200A company uses AWS Lake Formation to manage data lake permissions. The data lake contains sensitive customer data in the 'customer' database. The security team wants to ensure that only users with a specific tag 'access_level=analyst' can query the 'customer' table. Which combination of steps should the data engineer take to enforce this?
Medium201A company runs a time-series forecasting model that writes results to an S3 bucket every 5 minutes. A downstream ETL job reads this data, but sometimes fails because it encounters incomplete files (zero bytes). What is the MOST reliable way to ensure the ETL job only processes complete files?
Hard202A data engineering team needs to ingest streaming data from thousands of IoT devices and store it in Amazon S3 for batch processing. The data arrives at a rate of 10 MB/s, with occasional spikes up to 50 MB/s. The data must be processed in near real-time with minimal latency. Which AWS service should be used for ingestion?
Medium203A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for later analysis using Amazon Athena. The logs are generated continuously, and the total volume is about 1 TB per day. The data must be queryable within minutes of arrival. Which storage solution is most appropriate?
Easy204A data engineer manages an AWS Glue Data Catalog table that contains sensitive customer PII. The table's underlying data is in Amazon S3 and is queried by several AWS analytics services. The security team wants to implement column-level access control so that only authorized principals can view the PII columns, while other principals can still query non-sensitive columns. The solution must integrate with AWS Lake Formation and be enforced consistently across all query engines. Which approach should the data engineer take?
Medium205A financial services company is building a real-time fraud detection system. Transaction data is ingested via Amazon Kinesis Data Streams and processed by an Amazon Kinesis Data Analytics for Apache Flink application that runs sliding window aggregations. The output is written to an Amazon S3 bucket for downstream analysis. The Flink application is configured with parallelism of 4 and checkpointing every minute. The company has noticed that the application is experiencing high latency and the checkpointing is frequently failing. The CloudWatch metrics show that the Flink application's CPU utilization is near 100% and the checkpoint duration is spiking to over 5 minutes. The data engineer needs to improve performance. Which action should the data engineer take?
Hard206A company needs to protect sensitive data stored in Amazon S3 from unauthorized access. Which TWO actions should the data engineer take? (Choose two.)
Medium207A company uses Amazon DynamoDB as the primary data store for a high-traffic application. Recently, read latency has increased significantly. The DynamoDB table has on-demand capacity mode. Which action is MOST effective to reduce read latency?
Hard208A data engineer is monitoring an AWS Glue ETL job that intermittently fails with 'Container killed by YARN for exceeding memory limits' during a large shuffle stage. The job reads from Amazon S3, performs a groupByKey aggregation, and writes to Amazon S3. The engineer wants to reduce the chance of executor memory exhaustion without changing the source data. (Choose two.)
Medium209A company wants to grant cross-account access to an S3 bucket without using IAM roles. The data engineer needs to write a bucket policy that allows another AWS account to list objects. Which Principal should be specified in the bucket policy?
Medium210A data engineer manages an Amazon S3 data lake where analytics queries run through Amazon Athena. Monthly partition folders hold Parquet files, and each partition contains tens of thousands of small files averaging 40 KB. Athena queries that scan a single month take much longer than expected and consume far more bytes scanned than the actual data volume. The engineer must improve query performance without changing the table schema or the folder layout. What should the engineer do?
Medium211A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is ongoing, and the engineer notices that some transactions are not being replicated to the target. The DMS task is configured for full load plus change data capture (CDC). Which action should the engineer take to troubleshoot the missing transactions?
Hard212A data engineer is designing a data lake on Amazon S3 that must comply with a regulatory requirement to prevent any data from being overwritten or deleted for 7 years after creation. Which S3 feature should be used?
Hard213A data engineer is troubleshooting a Kinesis Data Streams consumer application that is falling behind. The stream has 10 shards and is receiving 5 MB/s of data. The consumer uses the Kinesis Client Library (KCL) with a single worker. The worker is processing all 10 shards but is experiencing high latency and checkpointing delays. Which THREE actions should the engineer take to improve consumer performance? (Select THREE.)
Hard214Refer to the exhibit. A data engineer checks the versioning status of an S3 bucket and sees the above output. The bucket contains critical logs that must not be permanently deleted. What should the engineer do to enhance protection against accidental or malicious deletion?
Easy215A data engineer is troubleshooting a Kinesis Data Streams application that is experiencing high latency. The stream has 2 shards. The application is using a single Kinesis Client Library (KCL) worker to process all shards. Which change will MOST likely reduce latency?
Hard216A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data is immutable and must be stored cost-effectively. Which AWS service should the engineer use?
Easy217A data engineer needs to transfer 50 TB of data from an on-premises data center to Amazon S3 over a 1 Gbps network. The transfer must be completed within one week. Which TWO AWS services can be used for this task? (Choose TWO.)
Easy218A data engineer manages an Amazon Redshift cluster that stores sales data in a table with a sort key on the sale_date column. The table is growing rapidly, and queries that filter by sale_date are becoming slower. The engineer notices that the table has a high percentage of unsorted rows. What should the engineer do to improve query performance with the least effort?
Medium219A company is using Amazon Kinesis Data Streams to ingest real-time clickstream data. The data is consumed by a fleet of EC2 instances running a custom application that processes the records and writes to DynamoDB. The application is experiencing high latency and records are being processed slower than they are produced. The stream has 5 shards. Which action would MOST effectively improve processing speed?
Hard220Which THREE of the following are valid storage classes in Amazon S3? (Choose THREE.)
Medium221A data engineer is designing a disaster recovery strategy for an Amazon RDS for PostgreSQL database. The primary database is in us-east-1. Which TWO approaches provide cross-region disaster recovery?
Medium222A data engineer needs to ensure that sensitive data stored in Amazon S3 is encrypted at rest. Which TWO options meet this requirement? (Choose TWO.)
Medium223A company is using Amazon ElastiCache for Redis to cache frequently accessed data. The cache hit ratio is low, and the engineering team suspects that the eviction policy is causing important data to be removed. Which eviction policy should be used to minimize eviction of the most frequently accessed keys?
Hard224A data engineer is designing an Amazon DynamoDB table for an order-processing application. The table uses a partition key of order_id and a sort key of order_date. The application needs to retrieve all orders for a specific customer within a date range, and the queries must be efficient at scale. The engineer must choose a design that supports these access patterns without full table scans. What should the engineer do?
Hard225A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query loads. The cluster uses a dc2.large node type with 2 nodes. Analysis shows that the workload involves frequent large table scans and complex joins. The engineer wants to improve query performance without changing the overall data volume. Which action should the engineer take?
Hard226Which THREE actions can help improve read performance in Amazon DynamoDB? (Choose THREE.)
Easy227A data engineer is ingesting streaming data from thousands of IoT devices into AWS. The data is JSON-formatted and must be stored in Amazon S3 for long-term analytics. Which service is most appropriate for real-time ingestion and routing to S3?
Easy228A data engineer is building a data lake on Amazon S3. The raw data arrives as JSON files, but the analytics team needs to query the data using standard SQL in Amazon Athena with optimal performance and minimal cost. The engineer wants to convert the JSON to a columnar format that supports predicate pushdown and is natively supported by Athena. Which storage format should the engineer choose?
Medium229A data engineer manages an Amazon S3 data lake with millions of small JSON files under prefixes partitioned by year/month/day. Amazon Athena queries scan far more data than expected and return slowly. The engineer wants to reduce bytes scanned and improve query performance while keeping files in S3 and queryable by Athena. Which solution meets these requirements with the LEAST operational overhead?
Medium230A healthcare company is ingesting patient data from a legacy system into an Amazon S3 data lake using AWS Glue. The legacy system produces CSV files with inconsistent schemas (columns may appear or disappear in different files). The data engineer needs to create a Glue ETL job that can handle schema evolution and transform the data into a standardized parquet format. The job should also be able to process new files as they arrive. Which approach should the data engineer use?
Hard231A data engineer is monitoring Amazon CloudWatch metrics for an Amazon Redshift cluster and notices high CPU utilization. The engineer wants to reduce CPU usage. Which TWO actions should the engineer take?
Easy232A company stores sensitive customer data in an Amazon S3 bucket with versioning enabled. A data engineer accidentally deleted the current version of an object. What is the quickest way to restore the object to its previous state without additional data transfer costs?
Hard233A data pipeline uses AWS Glue ETL jobs to process data from Amazon RDS for MySQL to Amazon S3. Recently, the jobs have been failing with the error 'Communications link failure' during the connection phase. The RDS instance is in a private subnet, and the Glue job uses a VPC endpoint for S3. What is the most likely cause?
Hard234A company uses AWS Glue to process sensitive data stored in Amazon S3. The security team requires that all data in transit between AWS Glue and S3 be encrypted. Which configuration should be used to meet this requirement?
Easy235An e-commerce company ingests clickstream data from their website into Amazon S3. The data is in JSON format, and each file is about 10 MB. They need to transform the data into a columnar format for analytics and load it into Amazon Redshift nightly. The transformation should be cost-effective and require minimal operational overhead. Which approach meets these requirements?
Medium236A data engineer is troubleshooting a failed AWS Glue job that reads from an Apache Hive metastore in an Amazon EMR cluster. The error message indicates 'ClassNotFoundException: org.apache.hadoop.hive.ql.metadata.HiveException'. The Glue job uses a custom Python shell script. What is the most likely cause of this error?
Hard237A data engineer maintains an AWS Glue job that reads JSON files from Amazon S3, applies a transform, and writes Parquet to a second bucket. The job's bookmark was enabled at creation, but each nightly run reprocesses all previously handled files, and downstream tables now contain duplicate rows. The job script has not been modified and the S3 prefix is unchanged. Which action will MOST directly resolve the duplicate processing?
Medium238A data engineer is tasked with transforming JSON data from an S3 bucket into Parquet format for efficient querying. The transformation should run on a schedule every hour. Which AWS service is best suited for this task?
Easy239A data engineer maintains an Amazon S3 data lake with millions of small JSON objects. The engineer needs to improve query performance by reducing the number of objects and compressing them into a columnar format that Amazon Athena can query efficiently. The data must remain partitioned by date. Which solution should the engineer use?
Medium240A data engineer notices that an AWS Glue ETL job processing data from Amazon S3 to Amazon Redshift has been failing intermittently with the error 'S3ServiceException: SlowDown'. Which action is MOST likely to resolve this issue?
Medium241An organization needs to audit all access to their S3 buckets for compliance purposes. They want to log both successful and failed API calls. Which AWS service should be used?
Medium242A company runs an Amazon Redshift cluster for analytics. During peak hours, query performance degrades significantly. The data engineer notices that disk space usage is above 80% on many nodes. Which of the following is the MOST effective long-term solution to improve query performance?
Hard243A data engineer needs to store semi-structured JSON event logs in a data lake on Amazon S3 and query them with Amazon Athena using SQL, including filtering on individual JSON attributes. The team wants to avoid transforming the files before querying. Which approach should the engineer use?
Easy244A company uses AWS Glue ETL jobs to process data from an S3 data lake. The job reads data in CSV format, transforms it, and writes to Parquet. The job runs daily and takes 2 hours to complete. The data volume is increasing by 20% each month. The engineer wants to reduce the job runtime. Which action is most effective?
Medium245A streaming application sends data to Amazon Kinesis Data Streams. The data must be enriched with reference data from an Amazon DynamoDB table in real-time. Which AWS service can be used to perform this enrichment with minimal latency?
Medium246A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer notices that the job fails with an error indicating that the Redshift table does not exist. The engineer has already created the Redshift cluster and database. What should the engineer do to resolve the error?
Hard247A company runs an Amazon RDS for PostgreSQL database and wants to capture change data (inserts, updates, deletes) to stream into Amazon Kinesis Data Streams for real-time processing. Which AWS service should be used to capture the changes directly from the database?
Easy248A company is designing a data lake on Amazon S3. The data includes CSV files, Parquet files, and images. The data engineering team needs to catalog the metadata and enable SQL queries. Which TWO AWS services should be used together?
Easy249A data engineer runs a Spark job on Amazon EMR that reads data from Amazon S3 and writes results back to S3. The job fails with an 'S3AccessDenied' error. The engineer verifies that the IAM role attached to the EMR cluster has s3:GetObject and s3:PutObject permissions on the relevant buckets. What is the MOST likely cause of the error?
Easy250A data engineer needs to ensure that an S3 bucket can only be accessed from a specific VPC. Which policy element should be used?
Medium251A data engineer receives an alert that a Kinesis Data Stream has a 'WriteProvisionedThroughputExceeded' error. The stream has 5 shards with 1 MB/s write capacity per shard. The producer application is sending data at 8 MB/s sustained. What should the engineer do to resolve the issue?
Easy252A data engineer needs to grant an IAM role read-only access to Amazon DynamoDB tables in a specific AWS account. Which IAM policy element should be used to restrict access to only the 'GetItem' and 'Query' actions?
Easy253An e-commerce company uses Amazon DynamoDB as the primary data store for its product catalog. The table has a simple primary key (ProductID) and handles 10,000 writes per second during peak hours. Recently, the engineering team noticed increased write latency and throttled requests during peak times. The table's provisioned write capacity is set to 12,000 WCU. What is the most likely cause of the throttling?
Medium254A data engineer manages an Amazon Redshift cluster that contains a table with credit card numbers. The security team requires that the credit card column be stored in encrypted form and that only users with a specific IAM role can see the full values. Other users should see a partially masked value when they query the table. Which Redshift feature should the data engineer use?
Hard255A data engineer needs to ingest data from an external partner's FTP server to Amazon S3. The data arrives once daily as a CSV file. Which AWS service should be used for this ingestion?
Easy256A data engineer is designing a solution to move data from an on-premises Oracle database to Amazon S3 using AWS DMS. The engineer needs to ensure that data changes are replicated continuously with minimal latency. Which DMS configuration is most appropriate?
Hard257A company is using Amazon Kinesis Data Firehose to ingest data into Amazon S3. The data must be transformed from JSON to Parquet format before delivery. Which feature should be enabled on the Firehose delivery stream?
Easy258A data engineer is configuring a data lake on Amazon S3 that contains sensitive customer information. The company requires that all access to this data be logged and monitored, and that any data shared with external partners must be anonymized before leaving the S3 bucket. Which combination of AWS services should the engineer use to meet these requirements? (Choose THREE.)
Medium259A media company ingests millions of small JSON files per day into an Amazon S3 bucket. Analysts run Amazon Athena queries over this data and report that each query scans far more data than the files matching their filters, resulting in high cost and slow performance. The files are partitioned by year/month/day in S3. What should a data engineer do to reduce the data scanned per query?
Hard260A data engineer needs to design a data ingestion pipeline that captures streaming data from mobile app events into Amazon S3 for analytics. The pipeline must support real-time processing of events and allow for schema evolution over time. Which AWS services should the engineer use? (Choose THREE.)
Medium261A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream and writes results to an S3 bucket. The application is consistently running out of memory and failing. The operator has already increased the Parallelism and TaskManager memory. What is the next BEST step to troubleshoot?
Hard262A company uses Amazon Redshift for its data warehouse. The data engineer notices that queries are slow on a large table that is frequently filtered on a column 'transaction_date'. Which optimization technique best improves query performance?
Hard263A data engineer uses AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe includes a 'Filter' step that removes rows where the 'status' column equals 'INVALID'. After running the recipe, the engineer notices that the output still contains rows with status 'INVALID'. The recipe was published and the job ran successfully. What is the most likely cause?
Hard264A company stores log files in Amazon S3. They want to automatically move logs older than 90 days to S3 Glacier Deep Archive to reduce costs. Which S3 feature should be used?
Medium265A data engineer needs to grant an AWS Glue ETL job access to read data from an Amazon S3 bucket that is encrypted with SSE-KMS using a customer managed key. The Glue job runs with an IAM role. Which action must the engineer take to allow the Glue job to decrypt the data?
Easy266A data engineer is designing a serverless data ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using AWS Lambda before being written to S3. Which two steps are required to enable this transformation? (Select TWO.)
Easy267A data engineer notices that an Amazon RDS for PostgreSQL instance's CPU utilization is consistently above 90% during business hours. The database is used for reporting queries. Which action should be taken FIRST to improve performance?
Easy268A company wants to grant read-only access to an S3 bucket for a data analyst. The analyst should be able to list objects and read object content. Which IAM policy effect and action combination is correct?
Easy269A data engineer needs to store semi-structured JSON data that is accessed infrequently but requires immediate retrieval when needed. The data must be durable and cost-effective. Which Amazon S3 storage class should be used?
Easy270A company is using AWS Lake Formation to manage permissions on a data lake. They want to grant a data scientist the ability to query tables in the 'analytics' database using Amazon Athena, but prevent them from accessing the underlying S3 data directly. What is the best way to achieve this?
Easy271A data engineer is using AWS Glue to run an ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job processes data in CSV format and the engineer wants to ensure the output is partitioned by year, month, and day based on a timestamp column in the data. The engineer needs to optimize the job for performance and cost. Which approach should the engineer take?
Medium272A data engineer has an AWS Glue job that processes data from an Amazon S3 bucket and writes to an Amazon Redshift cluster. The job is scheduled to run daily. Recently, the job started failing with the error: 'java.sql.SQLException: [Amazon](500310) Invalid operation: Spectrum Scan Error: S3 Access Denied'. The engineer verifies that the IAM role associated with the Glue job has full access to the S3 bucket. What is the most likely cause of this error?
Easy273A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. The table uses a partition key of deviceId and a sort key of timestamp, with a provisioned read capacity of 100 RCUs. During a sudden spike in traffic, the engineer observes throttling on read operations even though the consumed read capacity is well below the provisioned limit. What is the MOST likely cause of the throttling?
Medium274A data engineer is using AWS Glue Studio to create a visual ETL job that reads from an Amazon S3 bucket containing JSON files, applies a filter transformation, and writes the output to Amazon Redshift. The job must run daily. The engineer notices that the job is taking a long time to complete and wants to improve performance. Which action should the engineer take to optimize the job?
Medium275A healthcare company uses AWS Lake Formation to manage access to a data lake in Amazon S3. The data lake contains a table with patient records, and the company needs to ensure that only users in the 'Cardiology' department can query columns containing sensitive information such as patient name and diagnosis. Other departments should be able to query non-sensitive columns like patient ID and visit date. The company wants to implement this with the least operational overhead. What should the data engineer do?
Hard276A company uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The data contains personally identifiable information (PII) that must be redacted before storage. Which TWO actions can achieve this requirement? (Choose TWO.)
Hard277Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to decrypt objects in an S3 bucket using a customer managed KMS key. What is the effect of this policy?
Hard278A data engineer is using AWS Glue to run an ETL job that reads data from Amazon DynamoDB and writes to Amazon Redshift. The job fails with a 'ThroughputExceededException' error. What is the most likely cause?
Easy279A data engineer is troubleshooting a Kinesis Data Firehose delivery stream that is experiencing high error rates when writing to an S3 bucket. The error logs indicate 'AccessDenied' errors. The S3 bucket policy allows access from the Firehose service, but the errors persist. What is the most likely cause?
Hard280A company uses Amazon Redshift for data warehousing. The security team requires that all data stored in Redshift be encrypted at rest using a customer-managed KMS key. How should the data engineer configure this?
Medium281A data engineer is troubleshooting an AWS Glue ETL job that fails with the error: 'An error occurred while calling o123.pyWriteDynamicFrame. Access Denied when writing to S3 bucket: my-bucket'. The job uses a Glue service role named 'GlueServiceRole'. Which TWO actions should the engineer take to resolve the issue? (Choose TWO.)
Medium282A data engineer needs to store event data from IoT devices that arrives in bursts. The data is key-value and requires single-digit millisecond read and write latency. The engineer also needs to run complex analytical queries on the data for reporting. Which TWO services should be used together? (Choose TWO.)
Medium283A data engineering team needs to encrypt data at rest in an Amazon S3 bucket that stores sensitive customer information. The team must use an AWS Key Management Service (AWS KMS) customer managed key with automatic rotation enabled. Which configuration meets these requirements?
Medium284A media company ingests thousands of small JSON files per hour into an Amazon S3 bucket. A data engineer needs to convert these files into a compact, columnar format for efficient querying with Amazon Athena. The engineer wants to minimize storage costs and improve query performance. Which approach should the engineer take?
Hard285Which TWO options are valid ways to reduce storage costs for an Amazon S3 data lake that stores historical data rarely accessed after 30 days? (Choose TWO.)
Medium286A data engineer is using AWS Lake Formation to manage access to a data lake in Amazon S3. The engineer needs to grant a specific IAM role read access to only the columns 'customer_id' and 'purchase_amount' in a table stored in the AWS Glue Data Catalog. The table contains sensitive columns like 'credit_card_number'. Which Lake Formation permission model should the engineer use to achieve this?
Hard287A company uses Amazon RDS for MySQL to store transactional data. The database contains sensitive financial information. The company's security policy requires that all data at rest be encrypted using a customer-managed KMS key. The database was originally launched without encryption at rest. The security team now needs to enable encryption without significant downtime. What should they do?
Medium288A data engineer needs to monitor an AWS Glue ETL job that runs daily. The job sometimes fails due to missing partitions in the Data Catalog. The engineer wants to receive an alert when the job fails. What is the MOST operationally efficient way to achieve this?
Easy289A data engineer is using AWS Glue Studio to build a visual ETL job that reads JSON files from Amazon S3, applies a mapping transform, and writes Parquet to another S3 location. The engineer notices that the job is writing many small files, which hurts downstream query performance. Which action should the engineer take to reduce the number of output files without changing the source data?
Easy290A company stores sensitive data in Amazon S3. To meet compliance requirements, they need to ensure that any data older than 1 year is automatically moved to a lower-cost storage class. Which S3 feature should they use?
Easy291A data engineer is configuring an S3 bucket to host sensitive data. The security policy requires that all objects be encrypted with a key that is generated and managed by the customer, and that the key be stored in AWS KMS. Which encryption option should be used?
Medium292A data engineer is setting up a data pipeline to ingest streaming data from an IoT fleet. The data must be processed in near real-time and stored in Amazon S3 for analytics. Which THREE AWS services should the engineer consider using?
Easy293A data engineer is setting up a data pipeline that ingests streaming data from Amazon Kinesis Data Streams into an S3 data lake using Amazon Kinesis Data Firehose. The data contains personally identifiable information (PII). The security team requires that all data be encrypted at rest in S3 using an AWS KMS customer managed key (CMK) that is specific to the application. Additionally, the data must be encrypted in transit between all services. The engineer creates the KMS key and configures Firehose to use server-side encryption with the key for the S3 destination. However, Firehose delivery fails with an error indicating that the KMS key is not accessible. What is the most likely cause?
Medium294A data engineer is designing a data store for a real-time analytics application that requires low-latency reads and writes at scale. The data model includes time-series data with high ingest rates and queries that aggregate data over sliding time windows. The engineer needs a fully managed AWS service that supports automatic scaling and can handle millions of writes per second. Which service should the engineer choose?
Medium295A data engineer needs to schedule a recurring AWS Glue ETL job that must run every night at 02:00 UTC and must not start a new run while a previous run is still executing. The engineer wants the simplest managed scheduling option that integrates natively with Glue job run state. Which approach should the engineer use?
Easy296A data engineer uses Amazon EMR to run a Spark job that reads from S3 and writes to HDFS on the cluster. The job fails with an 'OutOfMemoryError: Java heap space' error in the executors. Which parameter adjustment should be made to resolve this?
Medium297Which TWO AWS services can be used to transform data in transit during ingestion? (Choose 2.)
Easy298A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to improve the performance of the Glue job, which currently takes several hours to complete. The job reads large Parquet files, performs joins, and writes to Redshift. Which two actions should the engineer take to improve performance? (Choose two.)
Hard299A company uses AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration completes successfully, but data validation shows some tables have missing rows. The task is configured for ongoing replication using change data capture (CDC). What is the MOST likely cause of the missing rows?
Hard300A media company stores video metadata in an Amazon DynamoDB table. The security team requires that all data at rest in the table be encrypted with a customer managed key in AWS KMS, and that the key usage be auditable. The data engineer needs to configure encryption for the table. Which action should the data engineer take?
Medium301A company's analytics team needs a petabyte-scale, fully managed data warehouse that supports standard SQL, columnar storage, and massively parallel query execution, and it must integrate with existing business intelligence tools with minimal operational effort. Which AWS service should the data engineer choose?
Easy302A company uses Amazon EMR to run Spark jobs on data stored in S3. After upgrading the EMR cluster to a new release, one of the Spark jobs fails with 'OutOfMemoryError' in the executor. Which configuration change is MOST likely to resolve this issue?
Medium303A data engineer runs an AWS Glue ETL job that reads CSV files from an Amazon S3 bucket, applies transformations, and writes Parquet output to another S3 bucket. The job fails with the error 'AnalysisException: Unable to infer schema for CSV. It must be specified manually.' The CSV files are stored with a header row, and the job's script uses the default Glue DynamicFrame reader without specifying format options. What is the MOST likely cause of the failure?
Medium304A data engineer needs to migrate an on-premises Apache Hadoop cluster to AWS. The cluster stores data in HDFS and runs MapReduce jobs. The company wants to minimize operational overhead and leverage serverless technologies where possible. Which AWS service should the data engineer use to replace HDFS storage?
Medium305A data engineer has set up an AWS Lambda function that processes files uploaded to an S3 bucket. The function is triggered by S3 event notifications. However, the function is not being invoked when a file is uploaded. The engineer checks the Lambda function's CloudWatch Logs and finds no execution logs. What should the engineer check FIRST?
Easy306A data engineer needs to transform CSV files arriving in an S3 bucket into Parquet format and store them in another S3 bucket. The transformation is simple and on-demand, triggered by data arrival. Which solution is the MOST cost-effective and requires the least operational overhead?
Medium307Order the steps to troubleshoot a failed AWS Glue job that reads from JDBC and writes to S3.
Medium308A data engineer notices that a nightly AWS Glue ETL job has been failing for the past three days with the error 'Unable to locate credentials'. The job uses an IAM role for execution. What is the most likely cause of this error?
Easy309A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer notices that some queries are waiting in the queue for a long time, and the WLM (Workload Management) configuration is set to auto. The engineer wants to implement manual WLM to improve query throughput and ensure that short-running queries are not blocked by long-running ones. Which TWO actions should the engineer take to achieve this? (Choose two.)
Hard310A data engineering team uses AWS Glue to extract, transform, and load (ETL) data from Amazon RDS for MySQL to Amazon S3. The job runs daily and processes incremental data. The team notices that the job is taking longer than expected. Which TWO actions can improve the job performance? (Choose two.)
Medium311A data engineer is configuring an AWS Lake Formation permissions model for a data lake in Amazon S3. Analysts must query a table through Amazon Athena and see only rows where the 'region' column equals 'EU'. The engineer has already registered the S3 location with Lake Formation and created the table in the AWS Glue Data Catalog. Which action should the engineer take to enforce the row-level restriction?
Hard312A data engineer stores application logs in an Amazon S3 bucket. Compliance requires that log objects be retained for seven years and that they cannot be deleted or overwritten by any user, including the account root user, during that period. The engineer must configure the bucket to enforce this. Which combination of settings should the engineer apply?
Medium313A company has a requirement to store audit logs for 7 years for compliance. The logs are stored in S3 and must be immutable. Which S3 feature should be used?
Hard314A company uses Amazon S3 to store sensitive financial data. The security team requires that all objects be encrypted at rest using AWS KMS with a customer-managed key. Additionally, they want to audit all KMS decrypt calls for compliance. Which configuration should be used to meet these requirements?
Hard315A company stores sensitive data in Amazon S3 and requires that all data in transit between on-premises applications and S3 be encrypted. The applications use the AWS SDK to upload and download objects. Which configuration should the data engineer implement to enforce encryption in transit?
Easy316A startup is building a mobile application that requires a database to store user profiles and preferences. The database must scale automatically with minimal administration. Which AWS service should they use?
Easy317Which TWO of the following are best practices for Amazon Redshift table design? (Choose TWO.)
Hard318A data engineer needs to schedule a daily AWS Glue job that extracts data from Amazon S3 and loads it into Amazon Redshift. The engineer wants to ensure the job runs at 2:00 AM UTC every day and can be monitored for failures. What is the simplest way to achieve this?
Easy319A data engineer is designing a pipeline that ingests JSON logs from an application into Amazon S3. The logs contain a timestamp field. The pipeline must partition the data by date in S3 (e.g., year=2024/month=10/day=01). Which approach minimizes transformation effort?
Medium320A company is migrating a large Oracle data warehouse to Amazon Redshift. Which THREE considerations are important for optimizing the Redshift cluster?
Hard321A data engineer needs to allow an AWS Lambda function to access a specific AWS KMS customer managed key to decrypt data. The Lambda function assumes an IAM role. Which policy statement should be added to the KMS key policy to grant the necessary permissions with least privilege?
Medium322A data engineer is building a data lake on Amazon S3 and needs to store JSON logs that will be queried by Amazon Athena. The engineer wants to minimize query cost and improve performance by reducing the amount of data scanned. The logs are approximately 1 KB each and arrive continuously. Which solution should the engineer implement?
Medium323A financial analytics company stores daily transaction records in Amazon S3 as Apache Parquet files, partitioned by year/month/day. The data engineering team queries these files with Amazon Athena. To reduce query runtime and cost, they want to apply fine-grained access control and column-level filtering without changing the files. Which solution should they use?
Medium324A data engineer is using AWS Step Functions to orchestrate an ETL workflow that includes an AWS Glue job. The Glue job occasionally fails due to transient issues, such as network timeouts. The engineer wants the Step Functions state machine to automatically retry the Glue job up to three times with exponential backoff before failing the workflow. Which Step Functions state configuration should the engineer use?
Medium325A data pipeline ingests streaming data from Kinesis Data Streams into S3 via Kinesis Data Firehose. Occasionally, small files are written to S3, increasing downstream processing costs. What is the most efficient way to reduce the number of small files?
Hard326A data engineer is building a near-real-time ingestion pipeline into Amazon S3. Small JSON files arrive continuously from thousands of devices, and the engineer must optimize the data lake for downstream Amazon Athena queries while minimizing storage cost and query latency. Which TWO actions should the engineer take? (Choose two.)
Hard327A company runs a data lake on AWS using S3 for storage and AWS Glue for ETL. The security team discovers that a contractor who left the company two months ago still has access to an S3 bucket containing sensitive data. The access was granted via an IAM user that was not deleted. The data engineer is asked to implement a solution to prevent future occurrences. The company uses AWS Organizations and has multiple accounts. The requirement is to automatically detect and remediate IAM users that have not been used for 90 days by disabling their access keys and notifying the security team. The solution must be least privilege and use AWS-native services. Which approach should the data engineer take?
Hard328A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is stored as CSV files and is updated daily. The engineer wants to load only new data each day without duplicating existing records. Which AWS service or feature should the engineer use to automate this process?
Easy329A data engineer must ingest a 4 TB Oracle database into Amazon S3 nightly. The database is on-premises and the network link supports only 200 Mbps. The engineer wants to minimize the total transfer time and avoid impacting production. Which approach should the engineer use?
Medium330An IAM role 'DataLakeRole' has the above S3 bucket policy attached to an S3 bucket. The role is assumed by an AWS Glue job. The Glue job is failing with 'Access Denied' errors when trying to list objects in the bucket. Which action should be added to the policy to fix the issue?
Hard331A company stores sensitive financial data in an Amazon Redshift cluster. The data engineer must ensure that all queries are logged for audit purposes and that the logs are stored in Amazon S3 with server-side encryption. Which THREE steps should the data engineer take to meet these requirements?
Hard332Refer to the exhibit. A data engineer deploys this CloudFormation template to create an AWS Glue job. The job fails on the first run with an error: 'AccessDeniedException: User: arn:aws:sts::123456789012:assumed-role/GlueServiceRole/... is not authorized to perform: s3:GetObject on resource: s3://my-bucket/scripts/etl.py'. What is the most likely cause?
Medium333A data engineer is running an Amazon EMR cluster with Spark to process log files. The cluster uses instance fleets with m5.xlarge core nodes. The engineer observes that the Spark job is running slower than expected. CloudWatch metrics show that the cluster's CPU utilization is below 20% but memory utilization is near 90%. Which configuration change would most likely improve performance?
Easy334A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow and the system is experiencing high disk usage. The engineer suspects that the distribution style is suboptimal. Which action should the engineer take to improve query performance?
Hard335A data engineer is using AWS Glue Studio to build a visual ETL job that joins a large Amazon S3 dataset with a small reference dataset of country codes. The join is currently implemented as a standard join, and the job runs slowly and shuffles large amounts of data. The engineer wants to optimize performance without changing the output. Which change should the engineer make?
Medium336A company uses Kinesis Data Analytics for SQL-based real-time analytics on streaming data. They notice that the application is processing data slower than the incoming rate, causing increased latency. Which action is MOST likely to improve the throughput?
Hard337A company uses AWS Glue to run ETL jobs that process data from Amazon RDS for MySQL and load it into Amazon S3. The job runs daily and processes incremental changes using the JDBC connection. Recently, the job has been failing with a 'Communications link failure' error. The RDS instance is in a private subnet. Which step should the engineer take first to diagnose the issue?
Hard338A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer needs to identify and resolve issues related to workload management (WLM). Which TWO actions should the engineer take to improve query performance? (Choose two.)
Hard339A data engineer is designing a multi-Region disaster recovery solution for an Amazon DynamoDB table. The table must be available in a secondary Region with minimal data loss and automatic failover. Which feature should be used?
Hard340A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket encrypted with SSE-KMS and writes to another S3 bucket also encrypted with SSE-KMS. The job uses an IAM role. The security team requires that the job have only the minimum necessary permissions to decrypt and encrypt data. Which TWO actions should be included in the IAM policy attached to the Glue job role? (Choose two.)
Medium341A company is building a data lake on Amazon S3. Data arrives from multiple sources in JSON, CSV, and Avro formats. The data must be transformed to Parquet and partitioned by date and source. Which TWO services can perform this transformation with minimal custom code? (Choose TWO.)
Medium342A social media company ingests user activity data from multiple sources using Amazon Kinesis Data Firehose. The data is delivered to Amazon S3 in near-real-time. The company wants to transform the data by adding a timestamp and masking email addresses before storing it in S3. The transformation should be applied to all records. What is the most cost-effective way to implement this transformation?
Medium343A company needs to ensure that data stored in Amazon RDS is encrypted at rest. Which action should the data engineer take?
Easy344A data engineer reviews the Glue job configuration. The job fails when processing large datasets. The error message indicates out-of-memory in the executors. Which change to the job configuration will most directly address this issue?
Hard345A company is using an Amazon RDS for MySQL database for its e-commerce platform. During a recent flash sale, the database experienced high read traffic, causing slow query performance. The company needs a solution that offloads read traffic with minimal application changes. Which action should be taken?
Medium346Which TWO actions can help optimize Amazon S3 storage costs for a data lake? (Choose two.)
Medium347A company uses Amazon RDS for MySQL to store financial data. A compliance requirement mandates that all database connections must be encrypted. Which configuration step is necessary?
Medium348A company uses Amazon S3 to store customer documents. The data engineer needs to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted with a customer-managed AWS KMS key. What should the data engineer do?
Easy349A company is migrating its on-premises data warehouse to Amazon Redshift. The data includes tables with up to 100 columns and 500 million rows. The migration involves a full load followed by incremental updates. The company needs to minimize downtime during the final cutover. Which THREE strategies should the data engineer use to facilitate the migration? (Choose THREE.)
Hard350A data engineer is designing a data pipeline that processes sensitive personal data. The data is ingested via Amazon Kinesis Data Firehose and stored in Amazon S3. The pipeline must ensure that the data is encrypted at rest and in transit. The engineer also needs to audit access to the data. Which combination of services meets these requirements?
Medium351A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format, and the engineer needs to convert it to Parquet before storage to optimize Athena queries. The Firehose delivery stream is configured with an AWS Lambda function for record transformation. However, the transformed data is still in JSON format in S3. What is the likely cause?
Medium352A data engineer is configuring an AWS Glue crawler to catalog data in an Amazon S3 bucket. The bucket contains CSV files organized in folders by year and month, and new files are added daily. The engineer wants the crawler to detect schema changes automatically and avoid reprocessing unchanged files on subsequent runs. (Choose two.)
Hard353A data engineer is building an AWS Glue ETL job that reads large Parquet datasets from Amazon S3 and must optimize performance and cost. The engineer wants to reduce the number of small files written to the target S3 prefix and improve read efficiency. (Choose two.)
Medium354A data engineer is troubleshooting an Amazon Redshift cluster that is running out of disk space. The engineer runs STV_PARTITIONS and notices that some slices have significantly more data than others. What is the most likely cause and solution?
Hard355A data engineering team needs to transform CSV files to Parquet format after they land in an S3 bucket. The transformation should be triggered automatically as soon as a new file arrives. Which AWS service is best suited for this task?
Easy356A company stores data in Amazon S3 with server-side encryption using AWS KMS (SSE-KMS). The data engineer needs to give a third-party auditor read-only access to the encrypted objects. The auditor has an AWS account. Which strategy should be used?
Hard357Which TWO actions can improve the performance of an AWS Glue ETL job that processes large datasets in Amazon S3? (Choose two.)
Medium358A data engineer needs to ensure that an Amazon Redshift cluster encrypts data at rest using a customer-managed AWS KMS key. Which configuration step is required?
Easy359A data engineer is configuring an AWS Glue ETL job that reads from and writes to an Amazon S3 bucket. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS, and that the job must fail if TLS is not used. Which two actions should the data engineer take to meet these requirements? (Choose two.)
Hard360A data engineer needs to audit data access in Amazon S3 for compliance. Which TWO services can be used to capture and analyze S3 access logs? (Choose TWO.)
Easy361A company uses Amazon DynamoDB to store user session data. The table has a partition key of UserID and a sort key of SessionStartTime. The application frequently queries for all sessions of a specific user within a date range. The table is provisioned with 1000 RCUs and 1000 WCUs. During peak hours, the application experiences throttling on read requests. Which action should a data engineer take to resolve the throttling with minimal changes?
Easy362A data engineer is preparing an AWS Glue ETL job that reads from and writes to Amazon S3 and must audit every access to sensitive data for compliance. The security team wants to know which principals accessed which objects and when, and also wants to detect anomalous access patterns. Which TWO AWS services should be used together to meet these requirements? (Choose two.)
Medium363A data engineer manages a large Amazon S3 data lake with millions of small JSON files ingested daily. Amazon Athena queries against this data lake are slow and costly due to high per-query data scanned. The engineer wants to optimize the storage layout to improve query performance and reduce cost, while keeping the data queryable in place. Which solution should the engineer implement?
Medium364A data engineer is configuring an AWS Glue crawler to catalog CSV files stored in Amazon S3. The files are organized in prefixes by year and month, and the engineer wants the crawler to detect new partitions automatically and avoid re-crawling unchanged partitions. (Choose two.)
Medium365A healthcare company stores patient records in an Amazon S3 bucket encrypted with SSE-KMS using a customer managed key. A new AWS Glue ETL job must read these records and write transformed data to another S3 bucket that is also encrypted with the same KMS key. The company's security policy requires that the Glue job's access to the KMS key be least-privilege and auditable. Which TWO actions should the data engineer take to meet these requirements? (Choose two.)
Hard366A data engineer needs to run an AWS Glue extract, transform, and load (ETL) job that joins an Amazon S3-based Parquet dataset with a slowly changing dimension table in Amazon Redshift. The Redshift cluster is in a private subnet and cannot be reached over the public internet. The engineer wants the Glue job to read from Redshift without exposing credentials in the job script. Which combination of actions should the engineer take to meet these requirements?
Medium367A data engineer must ensure that an AWS Glue ETL job can read from an Amazon S3 bucket encrypted with SSE-KMS and write to another S3 bucket also encrypted with SSE-KMS, using a single KMS key. The engineer has created an IAM role for the Glue job with permissions to access both buckets. What additional step is required to allow the Glue job to decrypt and encrypt data using the KMS key?
Medium368A company stores sensitive data in an S3 bucket. To meet compliance requirements, they must ensure that all objects are encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement should be applied to deny uploads that do not use the required encryption?
Medium369A data engineer manages an AWS Glue job that processes JSON files from Amazon S3 and writes Parquet to another S3 location. The job intermittently fails with 'Unable to find catalog table' errors, even though the table exists in the Glue Data Catalog. The job's IAM role has full S3 access but only limited Glue permissions. Which action will resolve the failure with the LEAST privilege?
Medium370A company is using Amazon S3 to store sensitive data. They need to ensure that all objects are encrypted at rest. Which combination of actions should be taken? (Choose TWO.)
Medium371A data engineer needs to store semi-structured JSON logs from multiple sources in a centralized data store for querying using SQL. The logs are immutable and need to be retained for 90 days. Which AWS service should be used?
Easy372A company is designing a data lake on Amazon S3. The security team requires granular access control based on data classifications. Which TWO AWS services can be used together to implement attribute-based access control (ABAC) for objects in S3?
Medium373A company has an Amazon S3 bucket with versioning enabled. They want to automatically delete noncurrent versions of objects after 30 days. Which lifecycle rule action should be used?
Medium374An e-commerce company uses AWS Glue to run ETL jobs that transform clickstream data from Amazon S3. The job reads Parquet files, performs aggregations, and writes the results to Amazon Redshift. The job runs successfully but takes longer than expected. The data volume is increasing. Which design change would MOST improve the job's performance?
Hard375An Amazon Kinesis Data Streams application is lagging behind. The data records are small (1 KB) and the shard count is 10. The consumer uses the KCL with default configuration. Which action will MOST effectively reduce the consumer lag?
Medium376A company wants to store data from thousands of IoT devices with varying data rates. The data must be stored in a schema-on-read fashion and support SQL queries. Which AWS service should be used?
Easy377A data engineer is building an AWS Glue job that reads from a large Parquet dataset in Amazon S3 partitioned by year/month/day and writes aggregated results to Amazon Redshift. The job currently reads all partitions and takes several hours. The engineer wants the job to process only partitions from the last seven days and reduce runtime. Which change should the engineer make?
Hard378Which TWO of the following are features of Amazon RDS Multi-AZ deployments? (Choose 2.)
Easy379A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. One of the Glue jobs occasionally fails due to transient issues, such as a temporary network glitch. The engineer wants the Step Functions state machine to automatically retry that specific Glue job up to three times before moving to a failure state. Which Step Functions feature should be used to implement this?
Medium380A data pipeline uses AWS Glue to process large CSV files. The team notices that some jobs fail with out-of-memory errors. Which TWO configuration changes can help mitigate this issue?
Hard381A data engineer is designing a data store for a real-time leaderboard application that requires sub-millisecond read and write latency. The leaderboard stores scores for millions of users and needs to be sorted by score. Which AWS service should the engineer use?
Medium382Which THREE steps are recommended for migrating an on-premises Oracle database to Amazon RDS for Oracle with minimal downtime? (Choose 3.)
Hard383A data engineer applies the bucket policy shown in the exhibit to an S3 bucket. The bucket contains sensitive data that must be encrypted at rest and accessed only over HTTPS. Which of the following statements is true?
Medium384A data engineer needs to keep a near-real-time copy of an Amazon DynamoDB table in Amazon S3 for analytics, capturing every item-level change with the before and after images and no impact on table write latency. Which approach meets these requirements with the LEAST operational effort?
Medium385A data engineer is designing a data ingestion pipeline to load millions of small JSON files from an on-premises FTP server into Amazon S3. The pipeline should minimize cost and operational overhead. Which approach is most suitable?
Medium386A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is ongoing, and the engineer needs to ensure that changes made to the source database during the migration are replicated to the target. The engineer has set up a full load plus change data capture (CDC) task. However, after the full load completes, the CDC task fails with an error indicating that it cannot find the archive log files. What should the engineer do to resolve this issue?
Medium387A data engineer is troubleshooting a daily batch ingestion pipeline that uses AWS Glue to read CSV files from Amazon S3 and write Parquet files to another S3 bucket. The job runs successfully but takes significantly longer than expected. The engineer notices that the input data is highly skewed with many small files. Which is the most effective optimization to reduce job duration?
Hard388A data engineer is troubleshooting a failed AWS Glue ETL job that reads from a JDBC source. The error log shows 'java.sql.SQLException: Connection timed out'. The job previously ran successfully. Which of the following is the MOST likely cause?
Hard389A data engineer must transform data in Amazon S3 using Apache Spark. The transformation logic needs to be reused across multiple AWS Glue jobs, and the engineer wants to version-control the code and run it in a serverless environment without managing clusters. Which approach should the engineer take?
Medium390A retail company uses Amazon Kinesis Data Firehose to ingest clickstream data from its website into an Amazon S3 bucket. The data includes fields: user_id, event_type, timestamp, page_url. Recently, the data engineering team noticed that some records have malformed JSON (missing commas, extra brackets) causing delivery failures to S3. The Firehose delivery stream is configured to retry failed records for 300 seconds, after which the records are sent to an S3 bucket for failed records. The team wants to transform the data to correct malformed JSON before delivery to the main S3 bucket. They need a solution that does not require managing servers and can handle high throughput. What should the team do?
Medium391Arrange the steps to create an AWS Glue job that transforms data from Amazon S3 to Amazon Redshift in the correct order.
Medium392A data engineer needs to ensure that an AWS Glue ETL job can access an Amazon S3 bucket that is encrypted with SSE-KMS. The Glue job runs with an IAM role. The KMS key policy grants access to the account root. Which TWO actions are required to allow the Glue job to read and write data in the bucket? (Choose two.)
Medium393A company stores its application logs in an Amazon S3 bucket. The logs are accessed frequently for the first 30 days, after which they are rarely accessed but must be retained for 7 years for compliance. The company wants to optimize storage costs while maintaining immediate retrieval availability for the first 30 days and the ability to retrieve logs within 12 hours after that. Which lifecycle policy should the data engineer configure?
Easy394A company uses AWS Glue to transform data in S3. The Glue job fails with memory errors. Which THREE actions can help resolve this?
Medium395A data engineer needs to grant an IAM user the ability to view Amazon CloudWatch Logs log groups and stream log events from a specific log group. Which IAM policy action should be used?
Easy396A data engineer is designing a real-time analytics solution using Amazon DynamoDB. The workload requires capturing all changes to a DynamoDB table and processing them in near-real-time to update a materialized view in Amazon Redshift. Which approach should the engineer use to capture and process the changes?
Hard397A data engineer is tasked with designing a disaster recovery solution for a data lake stored in Amazon S3. The data lake contains sensitive customer data that must be replicated to a different AWS Region. The engineer needs to ensure that all objects, including those with encryption using SSE-KMS, are replicated. Which solution meets the requirements?
Medium398A financial services company stores transaction records in an Amazon DynamoDB table. An audit requires that all data older than 7 years be automatically and permanently deleted. The data engineering team must implement this with minimal operational overhead and no application code changes. What should the team do?
Medium399A company needs to ingest real-time clickstream data from a web application into Amazon Redshift with minimal latency. The data volume is high and requires processing before loading. Which architecture is MOST appropriate?
Hard400A data engineer is configuring an AWS Glue crawler to catalog data stored in an Amazon S3 bucket. The data is partitioned by year, month, and day in a Hive-style structure (for example, s3://bucket/data/year=2023/month=01/day=15/). The engineer wants the crawler to recognize the partitions and add them to the AWS Glue Data Catalog. What should the engineer do?
Medium401A data engineer is building an Amazon Redshift data warehouse that ingests large staged files from Amazon S3 using the COPY command. The team wants to maximize load performance and minimize the time spent on ingestion. Which TWO practices should the engineer apply? (Choose two.)
Medium402A company stores sensitive data in an Amazon S3 bucket. To comply with regulations, all data must be encrypted at rest using server-side encryption. The security team wants to ensure that any attempt to upload an unencrypted object is automatically denied. Which S3 bucket policy condition should be used?
Medium403A data engineer is migrating an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB in size. The engineer needs to minimize downtime. Which AWS service should be used for the migration?
Easy404Match each AWS database service to its primary use case.
Medium405A team is designing a data lake on S3 and needs to enforce encryption at rest. They want to use server-side encryption with a KMS key that they manage. Which encryption option should they configure on the S3 bucket?
Medium406A company is using AWS Database Migration Service (DMS) to migrate a 2 TB MySQL database to Amazon Aurora MySQL. The migration is taking longer than expected. The source database is in a different AWS region. Which change would MOST likely improve the migration speed?
Hard407A company uses AWS CloudTrail to log all API calls. The security team wants to ensure that log files are tamper-proof and cannot be deleted. Which TWO actions should the data engineer take? (Choose TWO.)
Medium408A data engineer is building an AWS Glue Data Catalog table that references an Amazon S3 bucket containing CSV files. The security team requires that column-level access be restricted so that only specific IAM principals can view the column containing personally identifiable information (PII). The engineer needs to implement this restriction without modifying the underlying data. Which combination of actions should the engineer take?
Hard409A company needs to ingest data from a relational database into Amazon S3 for analytics. The database is an Amazon RDS MySQL instance. Which AWS service should be used for a one-time historical data load?
Easy410A data engineer is optimizing an Amazon Redshift cluster for a workload that includes frequent complex queries with multiple joins and aggregations. The engineer wants to improve query performance by using appropriate distribution styles and sort keys. Which TWO actions should the engineer take? (Choose two.)
Hard411A data engineer needs to share a dataset stored in an Amazon S3 bucket with another AWS account. The dataset must remain encrypted at rest using AWS KMS. The data engineer creates a bucket policy that grants the other account access to the bucket. However, the other account reports that objects appear encrypted and they cannot decrypt them. What is the most likely cause?
Hard412A company has an Amazon Redshift cluster with a mix of frequently accessed hot data and rarely accessed cold data. They want to reduce storage costs without affecting query performance for the hot data. Which strategy is MOST effective?
Hard413A data engineer is using AWS Glue ETL to transform data from an S3 data lake. The job fails with a memory error. Which approach should be used to resolve this issue without major code changes?
Medium414A company uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data delivery is delayed by up to 5 minutes. The engineer wants to reduce the delay to under 1 minute. Which parameter should be adjusted?
Easy415A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is then processed by a scheduled AWS Glue ETL job that loads it into an Amazon Redshift table. Recently, the Glue job has been failing with the error: 'S3ServiceException: Access Denied'. The Firehose delivery stream is configured with a prefix and error logging to the same S3 bucket. The Glue job uses the same IAM role that has s3:GetObject and s3:ListBucket permissions on the bucket. What is the most likely cause?
Medium416A data pipeline uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The delivery occasionally fails with 'Firehose is throttled'. What should be done to reduce throttling?
Easy417A data engineer is troubleshooting an AWS Glue job that fails with 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large dataset. Which TWO configuration changes should the engineer consider to resolve this issue? (Choose TWO.)
Medium418A data engineer is designing a pipeline that ingests streaming data into Amazon S3 using Amazon Kinesis Data Firehose. The data must be delivered to S3 in Parquet format and partitioned by date. The engineer needs to configure the Firehose delivery stream. Which two actions are required to meet these requirements? (Choose two.)
Medium419A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing sensitive customer records. The security team requires that the job's data be encrypted at rest using a customer managed AWS KMS key, and that the key policy restrict usage to the specific IAM role used by the Glue job. The engineer has already created the KMS key and attached the necessary IAM policy to the Glue job role. What additional step is required to ensure the Glue job can decrypt the S3 data using the customer managed key?
Medium420A data engineer is building an AWS Glue job that reads a table from the AWS Glue Data Catalog. The table contains columns with customer names, email addresses, and account numbers. The security team wants the job to mask the last four digits of account numbers in the output while leaving other columns unchanged. Which approach should the data engineer use?
Medium421A data engineer is designing an ingestion pipeline where AWS Lambda processes records from an Amazon Kinesis Data Stream. During peak traffic, records are being reprocessed and some are lost. The engineer needs to make the consumer resilient to failures and avoid duplicate processing. (Choose two.)
Hard422A data engineer is using AWS Lake Formation to manage permissions on a Data Catalog table backed by Amazon S3. Analysts query the table with Amazon Athena. The security team wants analysts to see only rows where the region column equals 'EU' and to prevent them from viewing the customer_id column entirely. Which combination of Lake Formation features should the engineer implement?
Hard423A company uses AWS Glue ETL jobs to transform data from Amazon RDS to Amazon S3 daily. The job recently started failing with memory errors. The data volume has grown 3x in the past month. Which change should the data engineer make to resolve the issue?
Medium424A data engineer is using AWS Lake Formation to manage permissions on a data lake in Amazon S3. The engineer grants SELECT permission on a table to an IAM role used by an Amazon Athena user. However, the user reports that queries against the table return an 'Access Denied' error. The engineer verifies that the IAM role has the necessary S3 permissions and that Lake Formation permissions are correctly set. What is the most likely cause of the error?
Hard425A data engineer needs to transform data in Amazon S3 using AWS Glue. The job must handle schema evolution and partition pruning. Which THREE features should be used?
Hard426A data engineer is designing a data ingestion pipeline for a social media analytics platform. The pipeline must ingest tweets in real-time, perform sentiment analysis, and store results in Amazon S3. The sentiment analysis is compute-intensive and must be done as the data arrives. The estimated throughput is 10,000 tweets per second. Which architecture is most suitable?
Hard427A data engineer receives an alert that an AWS KMS key has been scheduled for deletion by mistake. What is the immediate action to prevent the key from being deleted?
Easy428A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The workload includes frequent queries that retrieve all sessions for a user within a specific time range. The engineer notices that some queries are slow and wants to optimize the table design. Which action should the engineer take to improve query performance?
Hard429A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs and Amazon EMR steps as part of a nightly ETL pipeline. The pipeline occasionally fails due to transient issues such as Amazon S3 throttling or temporary network errors. The engineer wants to make the workflow more resilient without duplicating the entire state machine. Which Step Functions feature should be used to automatically retry failed states?
Hard430Refer to the exhibit. An IAM policy is attached to a user. What is the security implication of this policy?
Easy431A company uses AWS Glue to process streaming data from Amazon Kinesis Data Streams. The data is JSON formatted and includes a timestamp field. The company wants to partition the output in Amazon S3 by date and hour, and ensure exactly-once processing semantics. Which combination of configurations should be used?
Medium432A company uses Amazon Redshift for a data warehouse. They notice that queries are slow due to heavy data skew. Which optimization technique should be applied first?
Hard433A data engineer is designing a data lake on Amazon S3 to store JSON logs from an application. The logs are written once and never modified. The engineer needs to query the data using Amazon Athena with the best performance and lowest cost. The engineer wants to partition the data by year, month, and day based on the log timestamp. Which approach should the engineer use to organize the S3 objects?
Medium434A company uses AWS Glue to process sensitive data stored in S3. The security team requires that all data be encrypted at rest using customer-managed KMS keys. The data engineers are encountering 'Access Denied' errors when running Glue ETL jobs. What is the most likely cause?
Medium435A company is building a data lake on S3. They have a large volume of CSV files (hundreds of GB) in a source bucket. They need to convert them to Parquet, partition by date, and ensure the data is encrypted at rest with SSE-KMS. The pipeline must be triggered automatically when new files arrive. Which THREE steps should be part of the solution? (Choose THREE.)
Hard436A data engineer is monitoring an Amazon Redshift cluster and notices that the 'WLM query wait time' metric is consistently high during peak hours. The cluster uses automatic WLM. The engineer wants to reduce query wait times without changing the cluster size. Which action is MOST effective?
Hard437A data engineer is troubleshooting a DMS task that is replicating data from an on-premises Oracle database to an RDS for MySQL instance. The task is failing with 'ORA-1555: snapshot too old' error. What is the best course of action?
Hard438A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs that process data in Amazon S3. The engineer needs to ensure that if a Glue job fails, the entire workflow stops and an Amazon SNS notification is sent. Which Step Functions state type should be used to handle the error and send the notification?
Easy439A data engineer needs to transfer 10 TB of data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited to 100 Mbps, and the transfer must be completed within 48 hours. Which solution meets the requirements?
Medium440A data engineer manages an AWS Glue ETL job that reads CSV files from Amazon S3 and writes Parquet to another S3 prefix. The job recently started failing with the error 'AnalysisException: Unable to infer schema for CSV.' The engineer confirms the S3 path contains files and the IAM role has s3:GetObject permissions. The job's script uses glueContext.create_dynamic_frame.from_catalog with a database and table name. What is the MOST likely cause?
Medium441Which THREE of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose three.)
Hard442A data engineer needs to transform a large dataset stored in Amazon S3 using Apache Spark. The engineer wants to minimize startup time and use a serverless approach. Which AWS service should the engineer use?
Easy443A data engineering team needs to ingest streaming data from an application into Amazon S3 for analytics. The data volume is moderate and the team wants the lowest operational overhead. Which AWS service should they use?
Easy444A company wants to migrate its on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which AWS service should be used for the migration?
Easy445A data engineer is migrating an on-premises Apache HBase workload to Amazon DynamoDB. The HBase table has a row key with composite structure: customer_id (10 chars) + timestamp (10 digits). The access pattern is to query by customer_id and retrieve the latest entries. How should the DynamoDB table be designed to optimize performance?
Hard446Which TWO actions are effective ways to monitor the health of an Amazon DynamoDB table? (Choose two.)
Easy447A data engineer manages an AWS Glue Data Catalog used by Amazon Athena analysts. The security team wants column-level restrictions so that analysts querying a specific table cannot view the values in a cardholder_name column, while still being able to query all other columns. The analysts connect through Athena using an IAM role. Which approach meets this requirement with the LEAST operational overhead?
Medium448A data engineer manages an AWS Glue ETL job that processes sensitive customer records stored in Amazon S3. The security team mandates that all data at rest in the S3 bucket be encrypted with AWS KMS keys, and that the Glue job have the minimum permissions necessary to read and write data. The engineer creates a KMS key and configures the S3 bucket to use SSE-KMS with that key. Which additional step is required to allow the Glue job to access the encrypted data?
Medium449A marketing analytics team needs to ingest customer transaction data from an on-premises PostgreSQL database into Amazon S3 for analysis. The data volume is about 10 GB daily, and the team wants to perform full refresh daily (truncate and load) into S3 as Parquet files. The company has a Direct Connect connection to AWS. The team needs a simple, managed solution that minimizes operational overhead. What should the team use?
Easy450A data engineer is configuring an AWS Glue job bookmark on a job that reads partitioned Parquet data from Amazon S3 and writes to another S3 location. The engineer notices that reprocessing keeps occurring and wants the bookmark to correctly skip already-processed data. Which two actions should the engineer take? (Choose two.)
Medium451An organization wants to audit all API calls made to AWS services for compliance. Which AWS service should be used to capture and store these API calls?
Medium452A data engineer is reviewing an IAM policy that controls access to an S3 bucket. The policy is attached to a user group. The policy includes a condition that explicitly requires server-side encryption with SSE-S3 for all GetObject requests. The engineer notices that users are unable to download objects from the bucket. What is the likely cause?
Hard453A company has an Amazon Redshift cluster that stores petabytes of data. Queries are experiencing high disk usage due to large intermediate results. The data engineer needs to improve query performance without adding more nodes. Which action should the engineer take?
Hard454A data engineer is designing a data ingestion pipeline for IoT sensor data. The data is generated at a high velocity and must be processed in near real-time. The pipeline must also handle bursty traffic. Which TWO AWS services should be combined to achieve this? (Choose TWO.)
Medium455A company runs a real-time analytics platform using Amazon Kinesis Data Streams with a shard count of 10. The data is consumed by an AWS Lambda function that writes to an Amazon DynamoDB table. The DynamoDB table has a partition key of 'user_id' and a sort key of 'timestamp'. The table is provisioned with 5000 RCUs and 5000 WCUs. Recently, the application experienced increased write latency and throttling errors (ProvisionedThroughputExceededException) on the DynamoDB table. The CloudWatch metrics show that ConsumedWriteCapacityUnits averages 4500 with occasional spikes to 6000. The Lambda function’s concurrency is set to 1000. The data engineer suspects the issue is due to hot partitions. Upon investigation, the engineer finds that a small number of users generate a disproportionately large amount of data. Which course of action would best resolve the throttling while minimizing cost?
Hard456A data engineer is building an AWS Glue ETL job that must read from an Amazon S3 bucket in the same account and write to an Amazon Redshift cluster in a private VPC. The job must not traverse the public internet and must use least-privilege credentials. (Choose two.)
Hard457A company is using AWS Database Migration Service (DMS) to migrate a 2 TB Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is highly active with continuous writes. Which DMS migration type and additional configuration should the engineer use?
Medium458A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB in size. The migration must have minimal downtime and must be completed within a week. Which AWS service should the data engineer use to perform the migration?
Easy459A data engineer is building a pipeline that ingests records from an Amazon Kinesis data stream and writes them to Amazon S3 in Parquet format. The engineer wants to use AWS Glue to perform the transformation and needs the pipeline to handle records that arrive out of order and to deduplicate based on a record ID. Which combination of features should the engineer use?
Medium460A company runs a critical application on Amazon RDS for MySQL. To ensure high availability and automatic failover, the database is deployed as a Multi-AZ DB instance. The application uses read-heavy workloads. Which additional configuration should be used to offload read traffic without impacting write performance?
Medium461A company is using AWS Glue ETL to transform and load data from Amazon S3 to Amazon Redshift. The data engineer notices that the job is taking longer than expected. Which TWO actions can improve the job performance?
Medium462A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes 500 GB of data. The engineer notices that the job takes several hours and wants to optimize performance. The data is stored in Parquet format and partitioned by date. Which optimization should the engineer implement to improve the job's performance?
Hard463A company runs a data processing pipeline on Amazon EMR. The pipeline reads data from S3, processes it with Spark, and writes results back to S3. The engineer notices that the cluster is underutilized and wants to reduce costs. Which TWO actions should the engineer take? (Choose TWO.)
Medium464A data engineer uses AWS CloudTrail to investigate a security incident. The engineer runs the command shown in the exhibit. What does the output indicate?
Easy465A company needs to encrypt data in transit between an EC2 instance and an S3 bucket. Which method should be used?
Easy466A company needs to ingest data from multiple SaaS applications into Amazon S3. The data sources provide REST APIs. Which AWS service can be used to build a fully managed data ingestion pipeline without writing custom code?
Easy467A data engineer manages an Amazon Redshift provisioned cluster that serves a nightly ELT workload. The cluster's largest fact table is loaded with new rows each night, and queries frequently filter on a date column and join to a customer dimension. The engineer wants to improve query performance and reduce the time spent vacuuming. Which TWO actions should the engineer take? (Choose two.)
Hard468Refer to the exhibit. A data engineer ran the CLI command to check the configuration of an RDS instance named 'mydb'. Which statement accurately describes the current configuration?
Hard469A data engineer needs to run an AWS Glue ETL job that reads from an Amazon S3 bucket in another AWS account. The bucket owner has granted cross-account access, and the Glue job runs with an IAM role in the engineer's account. The job fails with an access denied error when reading the source objects. Which change is required to allow the Glue job to read the cross-account S3 data?
Easy470A data engineer is using AWS Step Functions to orchestrate an ETL workflow that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The workflow sometimes fails due to transient errors, and the engineer wants to implement a retry strategy that avoids duplicate data processing. Which TWO actions should the engineer take? (Choose two.)
Hard471Refer to the exhibit. A data engineer runs the AWS CLI command to describe a Glue job. The job is expected to process new data incrementally using job bookmarks. However, the job reprocesses all data every time it runs. What is the MOST likely reason?
Hard472Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to read objects from 'my-bucket' that have the tag 'classification=public'. The application account is 123456789012. However, the application is getting 'Access Denied' errors. What is the most likely reason?
Hard473A data engineer manages an AWS Glue ETL job that processes JSON files from an S3 bucket and writes Parquet to another bucket. The job uses a Glue DynamicFrame with a specified schema. During execution, the job fails with the error: 'AnalysisException: cannot resolve column 'transaction_id' given input columns: [txn_id, amount, timestamp]'. The source data has a column named 'txn_id', but the Glue job's script references 'transaction_id'. The job's catalog table for the source points to the correct S3 location and has the correct schema. What is the most likely cause of this error?
Medium474A data engineer needs to store a large volume of time-series data from IoT sensors. The data will be queried by timestamp and sensor ID, and the engineer wants to use a managed AWS database that can handle high write throughput and provide fast queries on recent data. The engineer also wants to automatically expire old data after 90 days to reduce storage costs. Which AWS service should the engineer use?
Easy475A company needs to ingest data from an on-premises Oracle database into Amazon S3 on a daily basis. The data volume is about 100 GB per day. Which AWS service is BEST suited for this task?
Easy476A data engineer is configuring an AWS Glue crawler to catalog data stored in Amazon S3. The data is organized as Parquet files under prefixes named by year, month, and day, such as s3://analytics/events/year=2024/month=05/day=17/. Queries in Amazon Athena must use partition pruning to limit scanned data. Which crawler configuration should the engineer choose?
Medium477A data engineering team is designing a data lake on Amazon S3 for storing sensor data from IoT devices. The data is written in near real-time and needs to be queried using Amazon Athena. Which TWO configurations should the team implement to optimize query performance and minimize costs?
Medium478A data engineer is using AWS Glue to run a nightly ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3 in Parquet format. The DynamoDB table is large and has a high read capacity. The engineer wants to minimize the impact on the DynamoDB table's performance and reduce the ETL job's runtime. Which approach should the engineer take?
Medium479Order the steps to migrate an on-premises database to Amazon RDS using AWS DMS.
Medium480A data engineer is building an AWS Glue ETL job that reads a large JDBC table from Amazon RDS for PostgreSQL. The job must read the table in parallel to reduce runtime, but the table has no numeric primary key or monotonically increasing column. Which AWS Glue connection property should the engineer configure to enable parallel reads?
Medium481A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a Redshift table. The data is in Parquet format and is partitioned by date in S3. The engineer wants to load only the data for the last 7 days to reduce load time and cost. Which Redshift command should the engineer use?
Medium482A company stores sensitive customer data in an S3 bucket. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key. However, when a data engineer attempts to upload an object using the AWS CLI, the upload fails with an access denied error. The engineer has s3:PutObject permission on the bucket. Which additional permission is most likely missing?
Medium483An e-commerce application uses Amazon ElastiCache for Redis to cache product catalog data. The cache currently uses lazy loading. The team wants to ensure that frequently accessed product data is always fresh. Which caching strategy should they implement?
Easy484A company uses AWS Glue to process JSON logs from S3. The logs have a nested structure and the schema evolves over time. The data engineer needs to ensure the Glue job can handle schema changes without failing. Which configuration should be used?
Hard485A company runs a SQL Server transactional database on Amazon RDS. They need to capture change data (inserts, updates, deletes) in near real-time and replicate them to an Amazon S3 data lake. Which AWS service is most suitable?
Medium486Which TWO practices improve the performance of AWS Glue ETL jobs? (Choose two.)
Medium487A company is using Amazon S3 to store sensitive customer data. The security team requires that all data be encrypted in transit and at rest. Additionally, they want to prevent any accidental public access. Which combination of actions should the data engineer take?
Hard488A company uses AWS KMS to encrypt data in Amazon S3. The security team wants to ensure that the KMS key can only be used from within the company's VPC. Which policy element should be added to the KMS key policy?
Easy489A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream, performs a 1-minute tumbling window aggregation, and writes results to an S3 bucket. Recently, the application started experiencing checkpoint failures and increasing processing delay. Which action should the engineer take FIRST to diagnose the issue?
Hard490A company is ingesting real-time financial transactions into Amazon Kinesis Data Streams. The data is then consumed by a Kinesis Data Analytics for Apache Flink application that calculates running totals. The application is experiencing high latency and checkpoint failures. Which TWO steps should the engineer take to improve performance and reliability? (Select TWO.)
Hard491A data engineer needs to schedule a daily AWS Glue ETL job that transforms data in Amazon S3. The job must run at 2:00 AM UTC every day. What is the simplest way to achieve this?
Easy492A data engineer is building an AWS Glue ETL job that reads JSON files from Amazon S3, flattens nested arrays, and writes Parquet to another S3 bucket. The job runs daily and processes about 2 TB. The engineer notices that job runs are failing intermittently with OutOfMemory errors during the shuffle phase. The job uses 10 G.1X workers. Which change should the engineer make to resolve the memory failures while minimizing cost?
Hard493A data engineer manages an Amazon Kinesis Data Stream with multiple shards. The stream is experiencing high throughput, and the engineer notices that some shards are throttling while others are underutilized. The engineer needs to redistribute the data evenly across shards to avoid throttling. Which action should the engineer take?
Hard494A data engineer is using AWS Step Functions to orchestrate a daily ingestion workflow. The workflow calls an AWS Glue job, then runs an AWS Lambda function to validate output, and finally starts an Amazon Redshift stored procedure. The engineer needs to ensure that if the Glue job fails, the workflow retries the Glue job up to three times with exponential backoff before failing the entire execution. Which Step Functions feature should the engineer configure?
Hard495A data engineer is building a streaming ingestion pipeline using Amazon Kinesis Data Streams. The producer application writes records with an explicit partition key derived from the device ID, and there are approximately 2,000 active devices. The engineer needs to ensure that records for the same device are processed in order by a downstream consumer. Which configuration should the engineer verify to guarantee per-device ordering?
Medium496A data engineer is configuring an Amazon S3 bucket for a data lake. The bucket must store sensitive financial data and comply with a regulation that requires all data to be encrypted at rest with keys that are automatically rotated every year. The engineer also needs to audit key usage and control access to the keys separately from other AWS services. Which encryption option should the engineer choose?
Hard497A data engineer is using AWS Glue to process a large dataset in Amazon S3. The dataset consists of many small JSON files (average 100 KB each) stored in a single prefix. The Glue job reads these files, performs transformations, and writes the output to Parquet in another S3 location. The job is running slowly and consuming many DPUs. Which action should the data engineer take to improve performance?
Hard498A company wants to ingest streaming data from thousands of IoT devices into Amazon S3 with minimal latency and then transform the data using Spark SQL. Which AWS service should be used for data ingestion?
Medium499A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it using Amazon Athena. The data is generated continuously and appended to S3 in small files. The engineer wants to optimize query performance and reduce costs. Which action should the engineer take?
Easy500A large e-commerce company uses Amazon DynamoDB to store shopping cart data. The table has a partition key of 'user_id' and a sort key of 'item_id'. The application performs frequent updates to the 'quantity' attribute for items in a user's cart. Recently, the operations team noticed that write requests are being throttled during peak shopping hours. The table is provisioned with 10,000 write capacity units (WCUs) and uses DynamoDB Accelerator (DAX) for read caching. The data engineer suspects that the throttling is due to hot partitions. The application uses a single AWS SDK client configured with retries. After reviewing the Amazon CloudWatch metrics, the engineer sees that the WriteThrottleEvents metric spikes for a few partition keys. The table has a high number of partitions. What should the data engineer do to resolve the throttling issue with minimal application changes?
Hard501A financial services company stores sensitive transaction data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys be automatically rotated every year. A data engineer needs to configure the S3 bucket to meet these requirements with minimal ongoing operational effort. Which solution should the engineer implement?
Hard502A data engineer needs to grant an AWS Lambda function permission to read objects from a specific Amazon S3 bucket. The Lambda function assumes an IAM role. Which policy should the engineer attach to the IAM role to allow the Lambda function to read objects from the bucket?
Easy503Which THREE are best practices for managing data in Amazon S3 for a data lake? (Choose three.)
Medium504Arrange the steps to set up a streaming ETL pipeline using Amazon Kinesis Data Firehose to Amazon S3.
Medium505A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains a column with inconsistent date formats and another with trailing whitespace. The engineer wants to apply these transformations reproducibly and schedule the recipe to run daily. Which combination of steps should the engineer take?
Hard506A data streaming application uses Kinesis Data Streams with 10 shards. The data producer is throttled frequently. Which action should be taken to resolve this issue?
Hard507Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket. What is the effect of this policy?
Hard508A company wants to implement least privilege access for its data lake on S3. Which THREE practices should be followed? (Choose THREE.)
Hard509A company uses Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed key that is rotated annually. Which encryption option should be used?
Easy510Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket named data-lake-bucket. The engineer wants to allow only GET requests from the corporate network (10.0.0.0/16) over HTTPS. However, users report that they cannot access objects even when connected to the corporate network. What is the issue?
Hard511A company is using Amazon RDS for MySQL and needs to automate backups with a retention period of 35 days. They also want to be able to restore to any point within the retention period. Which configuration should be used?
Medium512A data engineer needs to implement encryption for data at rest in an Amazon S3 bucket that stores sensitive financial records. The company's security policy requires that the encryption keys be managed by the company and rotated annually. The data engineer wants to use AWS Key Management Service (AWS KMS) to meet these requirements. Which solution should the data engineer implement?
Medium513A data engineer needs to monitor the number of records processed by an AWS Glue ETL job. Which CloudWatch metric should the engineer use?
Easy514A data engineer needs to load data from an Amazon DynamoDB table into an Amazon S3 bucket for analytics. The table is approximately 500 GB and has a high volume of write traffic. The engineer must minimize the impact on the table's read capacity and avoid consuming provisioned throughput. What is the MOST appropriate method to export the data?
Medium515A data engineer is designing a data lake on Amazon S3 that contains personally identifiable information (PII). The compliance team requires that all access to the data be logged and that any attempt to delete or modify data be detected and alerted. The engineer enables AWS CloudTrail data events for the S3 bucket and configures Amazon CloudWatch alarms. Which additional AWS service should the engineer use to automatically detect and remediate unauthorized changes to the S3 bucket's ACLs or policies?
Hard516A company is designing a data store for IoT sensor data that is written once and never updated. The data must be stored with high durability and low cost. Which TWO AWS storage services are most suitable? (Choose TWO.)
Medium517A data engineer is investigating intermittent failures in an AWS Step Functions state machine that orchestrates a nightly ETL workflow. The state machine invokes an AWS Glue job, then an Amazon EMR step, then an AWS Lambda function. Occasionally a task fails transiently and the entire workflow stops instead of retrying. The engineer needs the workflow to automatically retry failed tasks with exponential backoff before alerting. What should the engineer do?
Hard518A data engineer needs to grant an IAM role used by an AWS Lambda function permission to read encrypted data from an Amazon S3 bucket. The data is encrypted with a customer-managed AWS KMS key. The engineer wants to follow the principle of least privilege. Which combination of actions should the engineer include in the IAM policy for the Lambda execution role?
Hard519A company uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data is in JSON format and contains a 'timestamp' field with a Unix epoch value. The company wants to partition the S3 objects by year, month, day, and hour based on the timestamp. What is the MOST efficient method to achieve this?
Hard520A company wants to ingest real-time data from a social media API into Amazon S3 for analysis. The API provides data as JSON records. Which AWS service is best suited for this ingestion?
Easy521A data engineer is using AWS Glue to process data stored in Amazon S3. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS. The engineer has verified that the Glue job uses the AWS Glue Data Catalog and that the S3 bucket policy allows access. What should the engineer do to ensure that the Glue job enforces TLS when reading from and writing to S3?
Easy522A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer notices that queries are returning incorrect results, specifically missing some rows that are known to exist in the underlying data. The data is stored in Parquet format and is partitioned by date. The engineer runs a query with a WHERE clause on the date partition and finds that some dates are missing from the results. The S3 bucket contains folders for each date, but some folders are empty. What is the MOST likely cause of the missing rows?
Easy523A company needs to migrate an on-premises 10 TB PostgreSQL database to Amazon RDS for PostgreSQL with minimal downtime. Which AWS service should be used for the migration?
Easy524A company runs a data processing pipeline using Amazon EMR with Spark. The pipeline reads from S3, processes data, and writes to S3. Recently, the job started failing with 'S3AccessDeniedException' even though the EMR role has appropriate S3 permissions. Which TWO actions should the data engineer take to resolve this issue? (Choose TWO.)
Hard525A data engineer is using AWS Step Functions to orchestrate a complex ETL workflow that includes multiple AWS Glue jobs, Amazon EMR steps, and AWS Lambda functions. The engineer notices that on rare occasions, the entire workflow fails due to a transient error in one of the Lambda functions. The engineer wants to make the workflow more resilient without changing the overall architecture. Which approach is the MOST effective?
Medium526A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is in CSV format and is updated daily. The engineer wants to use a fully managed service that can handle the load without managing infrastructure. Which AWS service should the engineer use?
Easy527A company is using Amazon S3 to store log files. The security team requires that all data be encrypted in transit. Which of the following ensures encryption in transit for S3?
Easy528A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains inconsistent date formats in a column named 'transaction_date'. The engineer needs to standardize all dates to ISO 8601 format (YYYY-MM-DD) and then write the cleaned data to a new S3 location. Which transformation should the engineer apply?
Medium529A company stores its application logs in Amazon S3. The logs are generated daily and need to be retained for 3 years for compliance. The logs are accessed frequently for the first 30 days, occasionally for the next 6 months, and rarely after that. The data engineering team wants to minimize storage costs while ensuring that logs are available for retrieval within 12 hours for the first 6 months and within 48 hours after that. The team also wants to automatically delete logs after 3 years. Which lifecycle policy should the team implement?
Easy530A data engineer is using AWS Glue to run a job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job must be scheduled to run every day at 2:00 AM UTC. Which AWS Glue feature should the engineer use to set up this schedule?
Easy531A data engineer is using Amazon Redshift and needs to improve query performance for a workload that involves frequent joins between a large fact table and a small dimension table. The dimension table is updated daily with new records. The engineer wants to minimize data movement during joins. Which Redshift distribution style should the engineer use for the dimension table?
Hard532A data engineer is designing a data lake on Amazon S3. The compliance team requires that objects be automatically deleted after 7 years. Additionally, objects must be transitioned to Amazon S3 Glacier Instant Retrieval after 30 days to reduce costs. Which S3 lifecycle policy configuration meets these requirements?
Hard533A company wants to import data from an external FTP server into Amazon S3 on a daily basis. The data volumes are moderate. Which AWS service is MOST suitable for this task?
Easy534A data engineer is building a real-time data pipeline to ingest sensor data from IoT devices. The data is sent to AWS IoT Core, which publishes messages to a Kinesis Data Stream. Each message is about 1 KB in size. The data must be transformed (add a device location field) and then stored in Amazon S3 for long-term analytics. The engineer has set up a Lambda function to transform the records and write to S3. However, the engineer notices that the Lambda function is invoked thousands of times per second, causing high costs and occasional throttling. The Lambda function processes only one record at a time. The engineer wants to reduce the number of Lambda invocations and improve throughput. What should the engineer do?
Medium535Which TWO actions should a data engineer take to optimize Amazon S3 query performance for Amazon Athena when dealing with large Parquet files? (Choose 2.)
Medium536A data engineer needs to share a dataset stored in an S3 bucket with a partner AWS account. The partner should be able to read the data without needing to authenticate with the engineer's account. The engineer must not share any secret keys. Which approach should be used?
Hard537A data engineer needs to store semi-structured JSON data from IoT devices. The data is written once, read rarely, but must be queryable using SQL. The storage cost must be minimized. Which storage solution should the engineer choose?
Medium538Refer to the exhibit. A data engineer is reviewing the configuration of an Amazon Redshift cluster. The engineer wants to ensure that the cluster can be restored to a point in time up to 35 days in the past. Based on the exhibit, what change is needed?
Hard539A data engineer is configuring an Amazon Redshift cluster and needs to optimize query performance for complex analytical queries that involve large joins. The engineer wants to reduce the amount of data movement during query execution. Which two actions should the engineer take? (Choose two.)
Medium540Which THREE are valid considerations when troubleshooting data loss in an AWS Glue ETL job? (Choose three.)
Hard541A data engineer is setting up a data pipeline using Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using an AWS Lambda function before delivery. Which THREE steps are required to configure this?
Easy542A data engineering team is building a data lake on Amazon S3. They need to ingest data from multiple sources: (1) streaming IoT data, (2) daily CSV exports from an on-premises system via SFTP, and (3) change data capture (CDC) from an Amazon Aurora database. Which THREE services should the team use to ingest these data sources?
Hard543Refer to the exhibit. A data engineer has attached this IAM policy to an AWS Glue job role. The Glue job fails when trying to write transformed data to an S3 bucket located in a different AWS account. What is the most likely reason?
Medium544A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Aurora MySQL. After the migration, the data in Aurora is inconsistent with the source. The engineer needs to ensure ongoing replication with minimal downtime. Which solution should the engineer implement?
Medium545A data engineer needs to run a PySpark transformation on a 2 TB dataset stored in Amazon S3 and write the output back to S3 in Parquet. The team wants to use AWS Glue but does not want to manage clusters or tune Spark configuration manually. They also want to pay only for the time the job runs. Which AWS Glue component should the engineer use?
Easy546A company is using an Amazon RDS for PostgreSQL database to store application data. The data engineering team needs to run complex analytical queries that join multiple large tables. These queries are causing performance degradation on the production database. The team wants to offload the analytical workload to a separate system that can handle large-scale data processing. Which AWS service should the team use?
Easy547A company is using Amazon DynamoDB to store session data for a web application. The data engineer needs to ensure that the data is encrypted at rest. Which action should the data engineer take?
Medium548A data engineer is designing a data lake on Amazon S3 and needs to ensure that data is encrypted at rest. The company requires that encryption keys be managed by AWS and automatically rotated annually. The engineer also wants to audit key usage. Which S3 encryption option should the engineer choose?
Hard549A data engineer needs to audit all AWS KMS key usage events for the past 90 days to verify compliance. Which AWS service should be used?
Easy550A data engineer is designing a data ingestion pipeline using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format and must be converted to Apache Parquet before storage. The engineer wants to minimize costs and operational effort. Which two actions should the engineer take to meet these requirements? (Choose two.)
Medium551A data engineer is setting up cross-account access to an encrypted S3 bucket. The bucket uses a customer-managed KMS key. The engineer has configured the bucket policy and the IAM role in the source account. The target account still gets access denied errors when trying to read objects. What is the most likely cause?
Medium552A company stores application logs in Amazon S3 and needs to query them using standard SQL. The logs are in JSON format and are updated daily. The data engineering team wants a serverless solution that requires minimal management and can automatically discover the schema. Which AWS service should they use?
Easy553A data engineer is monitoring an AWS Glue job that reads from an Amazon S3 bucket and writes to Amazon Redshift. The job has been running for 2 hours, which is longer than usual. The engineer checks the Glue job's metrics and sees that the number of active executors is high, but the job is not making progress. The engineer suspects a data skew issue. Which action should the engineer take to diagnose and mitigate the skew?
Hard554A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to ensure that the job handles data quality issues such as duplicate records and missing values before loading. The job must also minimize the amount of data shuffled across the network. Which two actions should the engineer take? (Choose two.)
Medium555A data engineer stores raw customer records in an Amazon S3 bucket and runs an AWS Glue job that writes curated Parquet files to a second bucket. The governance team requires that the curated data carry a verifiable record of which job run produced it and that any modification to a curated file be detectable. The engineer must also prove that the curated dataset has not been altered since a nightly baseline. Which combination of AWS features should the engineer use?
Hard556A data engineer needs to store transaction data that requires strong consistency, ACID transactions, and complex join queries. Which AWS service is most appropriate?
Easy557A data engineer is using Amazon Kinesis Data Streams to ingest real-time data. The stream has 4 shards and is receiving 2 MB/s of data. The engineer notices that the WriteProvisionedThroughputExceeded metric is increasing. The engineer wants to resolve this issue with minimal changes. What should the engineer do?
Medium558A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format, and each record is approximately 5 KB. The company has set the buffer interval to 60 seconds and the buffer size to 5 MB. However, the data engineer observes that the delivery to S3 is delayed by up to 5 minutes during peak traffic. The engineer wants to reduce the delivery latency to under 1 minute. Which TWO actions should the engineer take? (Choose TWO.)
Medium559A data engineer applies the above IAM policy to a user. The user attempts to upload an object to the bucket 'my-data-lake' without specifying server-side encryption. What will happen?
Easy560A data engineer is troubleshooting an Amazon Redshift cluster that has been experiencing slow query performance. The engineer checks the system tables and finds that many queries are waiting on 'wlm_queued' time. The cluster has 10 nodes and uses automatic WLM. What is the most likely cause?
Hard561A data engineer runs an AWS Glue job that reads from an Amazon Kinesis Data Stream and writes to Amazon S3. The job must process records in order per shard and must checkpoint progress so it can resume after a failure without reprocessing all records. Which Glue configuration BEST supports this?
Hard562A company runs a data warehouse on Amazon Redshift. The data engineer notices that some queries are running slowly. Upon reviewing the system tables, the engineer finds that the 'svv_table_info' shows high 'unsorted' percentage for several large tables. What is the MOST effective action to improve query performance?
Hard563A data engineer is designing a data lake on Amazon S3 for a financial analytics workload. The raw data arrives as JSON files from an on-premises system. Analysts need to query the data using Amazon Athena with fast performance and minimal cost for queries that filter on a specific transaction date and customer ID. The engineer wants to convert the data to a columnar format that supports predicate pushdown and compression. Which storage format should the engineer choose?
Medium564A data engineer is configuring an AWS Glue ETL job that reads semi-structured JSON from Amazon S3 and must flatten nested arrays before loading into Amazon Redshift. During testing, the job fails with an AnalysisException stating that a column named 'events' cannot be resolved, even though the AWS Glue Data Catalog table shows the column. The engineer confirms the catalog table was created by a crawler and the S3 data is present. Which action will most directly resolve the schema resolution failure?
Medium565A data engineer manages an AWS Lake Formation governed data lake. Analysts query tables through Amazon Athena and must see only rows where the region column equals their assigned region, while column-level restrictions must also hide a national ID column. The engineer grants table SELECT to the analysts' IAM role in Lake Formation. What should the engineer configure next?
Hard566A data engineer is troubleshooting a slow-running Amazon Athena query. The query scans a large amount of data. Which TWO actions can improve query performance? (Choose TWO.)
Medium567A company is designing a data ingestion pipeline for clickstream data from a website. The data must be ingested in near real-time. Which TWO services can be used together to build this pipeline?
Medium568A company needs to share a dataset stored in an S3 bucket with a partner account. The dataset contains sensitive information, so the company wants to ensure that the partner account can only access the data using a specific VPC endpoint in the partner's account. Which S3 bucket policy condition key should be used?
Hard569A data engineer manages an Amazon Redshift cluster that experiences performance degradation during peak hours due to concurrent long-running queries and short ad-hoc queries competing for resources. The engineer wants to isolate the workloads so that short queries are not blocked by long-running ones, and to ensure that each workload gets a guaranteed share of memory and CPU. Which Redshift feature should the engineer implement?
Hard570A data engineer is troubleshooting a slow Amazon Redshift query. The query scans a large table with interleaved sort keys. The engineer notices that the query plan shows a sequential scan instead of a range-restricted scan. What is the MOST likely reason?
Hard571Refer to the exhibit. An AWS Glue ETL job is failing with an OutOfMemoryError. The job reads from Amazon S3 and performs a GROUP BY on a large dataset. Which change should the data engineer make to resolve this error?
Medium572A company uses S3 to store sensitive data. Which TWO S3 features can be used to protect data at rest?
Medium573A data engineer needs to load data from an Amazon DynamoDB table into an Amazon S3 data lake nightly. The table is large and the engineer wants to avoid consuming provisioned read capacity on the live table. Which approach should the engineer use?
Medium574A data engineer is using AWS Glue to read from an Amazon S3 bucket that contains data in Apache Parquet format, partitioned by year/month/day. The Glue job needs to read only the data for the last 7 days. The engineer wants to minimize the amount of data scanned and improve job performance. Which approach should be used to filter the partitions efficiently?
Hard575A data engineer needs to migrate an on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which approach should they use?
Medium576A company runs an e-commerce platform that generates clickstream data from user interactions on their website. The data is sent as JSON objects via HTTP POST to an API Gateway endpoint, which triggers a Lambda function that writes each record to a Kinesis Data Stream (100 shards). A second Lambda function consumes the stream, transforms the data (enriches with geolocation from a DynamoDB table), and writes to a Kinesis Data Firehose delivery stream that delivers Parquet files to an S3 data lake every 5 minutes. The system has been working for months, but recently the Firehose delivery stream started showing 'DeliveryFailed' errors for a subset of records. The errors point to 'InvalidData' from the Lambda transformation. The engineer reviews the Lambda transformation code and notices that the geolocation lookup occasionally fails because the DynamoDB table has a throttling issue. The engineer needs to handle these failures gracefully so that records that fail enrichment are still delivered to S3 with a null geolocation field, without blocking other records. Which course of action should the engineer take?
Hard577A company uses DynamoDB with global tables in two AWS Regions. The data engineer observes that a write to the table in us-east-1 is not immediately visible in a read from eu-west-1. What is the most likely reason?
Hard578A data engineer is using AWS Glue job bookmarks to process incremental data from Amazon S3. The job reads from a partitioned S3 path and writes to Amazon Redshift. After a recent run, the engineer notices that some new partitions were not processed. The job bookmark state shows that the job has already processed up to a certain timestamp. What is the most likely reason for the missing partitions?
Hard579Refer to the exhibit. A data engineer is using a Kinesis Data Stream with one shard. The application writes 2000 records per second, each 1 KB. The put record calls are frequently throttled. What is the most likely cause?
Medium580A data engineer is using AWS Glue to run a PySpark ETL job that processes millions of small JSON files stored in an Amazon S3 bucket. The job is experiencing high memory usage and failing with 'Container killed by YARN for exceeding memory limits'. The engineer wants to optimize the job to handle the data more efficiently without changing the source data format. Which solution will MOST effectively reduce memory usage and improve performance?
Medium581A company uses Amazon Redshift for analytics. They notice that some queries are slow due to data redistribution. The data engineer wants to minimize data movement across nodes. Which table design strategy should be used? (Choose TWO.)
Hard582Order the steps to set up an Amazon EMR cluster for processing data in S3 using Spark.
Medium583A company uses AWS Lambda to process messages from an Amazon SQS queue. The messages contain JSON payloads that need to be transformed and written to an Amazon DynamoDB table. Recently, the Lambda function has been timing out and messages are being sent to the dead-letter queue (DLQ). What is the BEST way to troubleshoot and resolve this issue?
Medium584A data engineer is troubleshooting an AWS Glue job that writes data to an Amazon S3 bucket in Parquet format. The job runs successfully but the output files are smaller than the configured 'groupFiles' size. The engineer has set 'groupFiles' to 'inPartition' and 'groupSize' to 1 GB. The input data is 10 GB in a single partition. What is the most likely reason for the small files?
Hard585A company is migrating its on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is 2 TB and runs on a single server. Which AWS service should be used for the migration?
Medium586A data engineer manages an AWS Glue Data Catalog shared across teams. Analysts in one team must be able to query only the sales database and its tables, while another team owns the marketing database. The engineer wants permissions managed centrally in Lake Formation and wants the analysts to be able to create their own temporary tables but not alter the sales tables. Which combination of Lake Formation grants should the engineer apply?
Hard587A company uses AWS Glue to run ETL jobs daily. The data engineer wants to reduce costs by optimizing the job configuration. Which two actions will help reduce costs? (Choose TWO.)
Easy588A data engineer sees the CloudWatch log entry in the exhibit for a Lambda function that processes data from an Amazon SQS queue. What is the MOST likely cause of the timeout?
Medium589A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an on-premises Oracle database to an Amazon S3 bucket in near real time. The source table has a primary key and the database is configured for ARCHIVELOG mode. The engineer needs change data capture (CDC) to capture INSERT, UPDATE, and DELETE operations. Which AWS DMS task setting should be used to capture ongoing changes?
Medium590A data engineer wants to ensure that only users with a specific tag (e.g., "Department": "DataEngineering") can access an S3 bucket. How can this be enforced?
Easy591A data engineer is setting up an Amazon Redshift cluster and needs to load data from Amazon S3. The data is in CSV format and contains a large number of rows. The engineer wants to achieve the fastest possible load time. Which method should the engineer use?
Easy592Refer to the exhibit. A data engineer runs the command on an object in S3. The engineer expected the object to have a tag 'type=raw' but sees no metadata. What is the likely cause?
Hard593A data engineer is designing a data ingestion pipeline for IoT sensor data. The sensors send JSON messages every second. The data must be available in Amazon S3 within 5 minutes and must be transformed (JSON to Parquet) before storage. Which combination of services meets these requirements?
Hard594A company uses Amazon RDS for PostgreSQL with encryption at rest using AWS KMS. The company needs to share a database snapshot with a different AWS account. What must be done to allow the target account to restore the snapshot?
Hard595A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another. The security team requires that all data be encrypted at rest and that the job use a customer-managed KMS key for encryption. The engineer configures the job with `--encryption-type sse-kms` and a KMS key ID. However, the job fails with an access denied error when writing to S3. The IAM role used by the Glue job has permissions to read and write the S3 buckets but has no KMS permissions. Which additional IAM permissions are required for the Glue job role to successfully write encrypted data?
Medium596A data engineer maintains an Amazon DynamoDB table that stores device telemetry. The table uses a partition key of deviceId and a sort key of timestamp, with on-demand capacity mode. A new fleet of devices writes data with a deviceId pattern that hashes to a small number of partitions, and the engineer observes throttling on writes even though consumed capacity is well below any configured limit. Which change addresses the root cause?
Hard597A data engineer is designing a data lake on Amazon S3 and needs to catalog data using the AWS Glue Data Catalog. The data is stored in Parquet format, partitioned by year/month/day. The engineer wants to query the data using Amazon Athena and ensure that partition pruning occurs to minimize query costs. Which action should the engineer take?
Hard598A data engineer is building a streaming pipeline using Amazon Kinesis Data Streams and AWS Lambda. The Lambda function processes records and writes to Amazon DynamoDB. The engineer notices that the Lambda function is throttled during high traffic. Which action should the engineer take to reduce throttling?
Hard599A data engineer is troubleshooting an AWS Glue ETL job that suddenly started failing with 'An error occurred while calling o103.pyWriteDynamicFrame. Unknown error'. The job writes data to an Amazon Redshift table. Which step should the engineer take FIRST?
Hard600A company needs to ingest streaming data from multiple sources and store it in Amazon S3. The data volume is up to 5 GB per hour. What is the MOST cost-effective ingestion service?
Easy601A data engineer is responsible for an Amazon Redshift cluster that stores financial data. The security team requires that all connections to the cluster from outside the VPC use SSL, and that the cluster's audit logs capture connection and user activity. The engineer has already enabled audit logging to Amazon S3. Which additional configuration should the engineer apply to meet the SSL requirement?
Medium602A data engineer runs the AWS CLI command to retrieve the lifecycle configuration of the 'my-data-lake' bucket. The output is shown in the exhibit. What is the effect of this lifecycle policy?
Medium603A data engineer is optimizing an Amazon S3 data lake that stores large volumes of JSON logs. The engineer wants to reduce storage costs and improve query performance in Amazon Athena. Which TWO actions should the engineer take? (Choose two.)
Medium604A data engineer needs to grant an IAM user access to query a specific table in Amazon Athena, but the user should not be able to view other tables in the same database. Which method should the engineer use?
Easy605A data engineer must load a 250 GB uncompressed CSV dataset from Amazon S3 into Amazon Redshift. The data must be loaded daily, and the engineer wants to minimize load time and avoid saturating the cluster's leader node. Which approach meets these requirements?
Easy606A data pipeline uses AWS Glue to process data from an S3 data lake. The pipeline fails intermittently with a 'ThrottlingException' when writing to a DynamoDB table. What is the MOST likely cause?
Easy607A company stores IoT sensor data in S3 as JSON files. They need to convert the data to Parquet format for efficient querying with Amazon Athena. Which AWS service can perform this transformation with minimal effort?
Easy608A data engineer is building an Amazon DynamoDB table that will store IoT sensor readings. Each reading has a deviceId (partition key) and a timestamp (sort key). The team wants to retrieve all readings for a device within the last 24 hours, and they also want to minimize the number of read capacity units consumed. Which access pattern should the engineer implement?
Hard609A data engineer is using AWS Lake Formation to manage access to a data lake stored in Amazon S3. The engineer grants SELECT permission on a table in the AWS Glue Data Catalog to an IAM role used by an Amazon Athena user. However, the user still cannot query the table and receives an 'Access Denied' error. The engineer verifies that the IAM role has the necessary AWS Lake Formation permissions and that the S3 bucket policy allows access. What is the most likely cause of the issue?
Hard610A company is using AWS Glue to process streaming data from Amazon Kinesis Data Streams. The job fails intermittently with a 'MemoryError' when the stream has a sudden spike in data volume. Which configuration change would best prevent this error?
Medium611A company uses AWS Glue ETL jobs to transform data in Amazon S3. The data arrives in JSON format but needs to be converted to Parquet for efficient querying. Which AWS Glue feature should be used to infer the schema and generate transformation code?
Easy612A company wants to transform data in Amazon S3 using SQL queries without provisioning servers. The transformations are ad-hoc and run occasionally. Which service should be used?
Easy613A data engineer is setting up an AWS Glue job that reads from an Amazon Kinesis Data Stream and writes to an Amazon S3 bucket. The security team requires that all data in transit be encrypted using TLS, and that the Glue job must use a VPC endpoint to access Kinesis and S3. Which configuration ensures compliance?
Medium614Refer to the exhibit. A data engineer runs the above CLI command and sees the output. The security team requires that the RDS instance not be accessible from the internet. Which change should the engineer make?
Medium615A data engineer is designing a data lake on Amazon S3 and needs to ensure that objects are automatically encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement achieves this?
Medium616A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. Each item has a partition key of deviceId and a sort key of timestamp. The table is configured with on-demand capacity mode. The engineer needs to retrieve all readings for a specific device within the last 24 hours, and the query must return results sorted by timestamp in ascending order. Which operation should the engineer use?
Medium617A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The engineer needs to retrieve all sessions for a specific user that started within the last 30 days. Which operation should the engineer use to achieve the lowest latency?
Hard618An e-commerce company wants to capture clickstream data from its website and store it in Amazon S3 for analytics. The data arrives continuously and the company needs near-real-time processing. Which solution is most appropriate?
Easy619A data engineer maintains an Amazon Kinesis Data Firehose delivery stream that writes JSON records to Amazon S3 and then invokes an AWS Lambda function for transformation. The Lambda function occasionally times out, causing records to be delivered untransformed. The engineer must ensure failed records are captured for later reprocessing without blocking delivery. What should the engineer do?
Medium620Refer to the exhibit. A data engineer is configuring an IAM policy for a Lambda function that writes transformed data to S3. The function writes to both 'example-bucket/data/' and 'example-bucket/public/'. The policy is intended to enforce server-side encryption with SSE-S3 for all objects written to the 'public/' prefix, while allowing all operations on other prefixes. However, the Lambda function is failing with an AccessDenied error when writing to 'example-bucket/public/'. What is the most likely cause?
Hard621A company is using an Amazon RDS for MySQL database for an e-commerce application. During a sales event, the database experiences high read traffic, causing slow query performance. The company wants to reduce the read load on the primary database without changing the application code. Which solution meets these requirements?
Medium622A data engineer is designing a data ingestion pipeline for real-time clickstream data. Which TWO services can be used to ingest the data into Amazon Kinesis Data Streams?
Easy623A data engineer is building an AWS Lambda function that processes records from an Amazon Kinesis data stream. The Lambda function needs to read from the stream and write processed data to an Amazon S3 bucket. The security team requires that all data in transit be encrypted using TLS, and that the Lambda function authenticate to Kinesis and S3 using temporary credentials. Which combination of configurations should the engineer use?
Medium624A company has multiple AWS accounts and wants to centrally manage permissions and access to data lakes. They have enabled AWS Organizations and want to use a single set of policies that apply to all accounts. Which policy type should be used at the organization level?
Hard625Which TWO actions should a data engineer take to protect sensitive data in an Amazon S3 bucket from being accessed by unauthorized users? (Select TWO.)
Medium626A data engineer is setting up an Amazon S3 bucket to store large CSV files that will be queried using Amazon Athena. The engineer wants to minimize query costs and improve performance. The files are currently stored in a single prefix without any partitioning. The most common queries filter data by `year` and `month`. What should the engineer do to optimize the Athena queries?
Easy627Refer to the exhibit. A data engineer is troubleshooting a Kinesis Data Streams consumer that is falling behind. The stream has 2 shards and is receiving data at a rate of 2 MB/s. The consumer is an AWS Lambda function with a batch size of 100 records. What should the engineer do to improve consumer throughput?
Medium628A data engineer needs to transform JSON data from an S3 bucket into Parquet format and load it into Amazon Redshift. The transformation must be performed incrementally as new data arrives. Which AWS service is BEST suited for this task?
Easy629A data engineer is using Amazon Managed Workflows for Apache Airflow (MWAA) to orchestrate a data pipeline. The pipeline includes a task that runs an AWS Glue job. The engineer notices that the Glue job occasionally fails due to transient issues, and the Airflow task fails immediately without retrying. The engineer wants to configure the Airflow task to retry the Glue job up to 2 times with a 5-minute delay between retries. Which configuration in the Airflow DAG should the engineer use?
Hard630A company is ingesting streaming data from multiple sources using Amazon Kinesis Data Streams. The data is then processed by an AWS Lambda function that transforms the records and writes them to an Amazon S3 bucket. The Lambda function is failing intermittently with timeout errors. The average record size is 5 KB, and the shard count is 2. What is the MOST likely cause of the timeout errors?
Hard631A data engineer is troubleshooting a slow-running query on an Amazon Redshift cluster. The query involves joining two large tables. The engineer notices that the query plan shows a large number of distribution and broadcast operations. Which design change would most likely improve query performance?
Medium632Refer to the exhibit. A data engineer runs this AWS Glue Data Catalog DDL statement to create a table. The CSV files in 's3://my-bucket/sales/' use a pipe delimiter (|) instead of a comma. What change is needed to correctly read the data?
Easy633A data engineer needs to transform JSON data into CSV format using AWS Glue. The transformation is simple and must be executed on a schedule. Which Glue component is MOST suitable?
Easy634An e-commerce company is building a near-real-time dashboard to monitor customer clickstream data. The data is ingested via Amazon Kinesis Data Streams, transformed using AWS Lambda, and stored in Amazon S3. The team needs to query the data using Amazon Athena. Which THREE steps should be taken to optimize cost and performance? (Choose three.)
Medium635A data engineer must load data from an Amazon S3 bucket into an Amazon Redshift cluster as part of a nightly batch pipeline. The source files are already in Parquet format and include columns that map directly to the target table. The engineer wants the fastest, most cost-effective load method that avoids staging the data through an external service. Which command should the engineer use?
Easy636A small startup is building a data pipeline to ingest customer orders from a web application into Amazon Redshift for analytics. The orders are written to an Amazon RDS MySQL database. The startup wants to replicate the orders to Redshift in near-real time (within 5 minutes) with minimal operational overhead. The data volume is low, averaging 100 new orders per minute. The startup has a single data engineer who is also responsible for other tasks. What is the simplest solution?
Easy637A data engineer needs to securely store database credentials for an RDS instance. Which TWO AWS services can be used?
Easy638Match each AWS Glue component to its role.
Medium639A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an Amazon RDS for MySQL database to an Amazon S3 bucket in Parquet format. The replication task is configured with full load plus change data capture (CDC). After several hours, the engineer notices that the S3 bucket contains only the full load data and no incremental changes. Which action should the engineer take to ensure CDC changes are captured?
Medium640A company uses AWS Glue to catalog data in Amazon S3. The data includes personally identifiable information (PII). The security team requires that PII be masked when queried by users who are not data owners. Which AWS service should be used to enforce this requirement?
Medium641A data engineer is responsible for a data pipeline that uses Amazon S3 as a data lake, AWS Glue for ETL, and Amazon Athena for ad-hoc queries. The pipeline ingests CSV files from an external partner via SFTP into an S3 bucket. The files are then processed by a Glue job that converts them to Parquet and writes to a separate S3 bucket partitioned by date. The Glue job runs daily and is triggered by a scheduled CloudWatch Events rule. Recently, the data engineer noticed that some days the Glue job fails because of memory errors, and on those days the Athena queries that rely on the data return incomplete results. The engineer needs to ensure that the pipeline is resilient and that Athena queries always see a complete view of the data, even if the Glue job fails mid-run. The engineer also needs to minimize re-processing of data. Which course of action should the engineer take?
Hard642A data engineer is optimizing an AWS Glue ETL job that processes large Parquet files in Amazon S3. The job currently takes several hours to complete. The engineer wants to improve performance by tuning the job's execution parameters. Which TWO actions will MOST effectively reduce the job's runtime? (Choose two.)
Hard643A data engineer is configuring cross-account access so that an analytics AWS account can read objects from a data-lake S3 bucket in a producer account. The objects are encrypted with SSE-KMS using a customer managed key in the producer account. The engineer has already added a bucket policy granting s3:GetObject to the analytics account's IAM role. Reads still fail with AccessDenied. Which additional change is required?
Medium644A data engineer is designing a data lake on Amazon S3 and needs to store data in a format that supports schema evolution and efficient columnar storage. The data will be queried using Amazon Athena and Amazon Redshift Spectrum. The engineer wants to minimize storage costs and improve query performance. Which storage format should the engineer choose?
Medium645A company uses Amazon S3 to store historical stock market data as CSV files. They run daily Amazon Athena queries to generate reports. Recently, the finance team reported that queries are timing out and costs have increased significantly. The data engineering team notices that the S3 bucket contains thousands of small files (average 100 KB) due to a misconfigured ingestion pipeline. They need to improve query performance and reduce costs without changing the existing reporting schedule. The team has access to AWS Glue and can create new tables. Which solution should they implement?
Medium646A company is migrating an on-premises Hadoop cluster to AWS. The cluster processes large files in CSV format using Apache Spark. Which data store should be used as the primary storage for the data lake to optimize cost and performance?
Medium647A financial services company stores transactional records in Amazon DynamoDB. Auditors require that any item be recoverable to its exact state from any point within the last 30 days, including after an accidental delete or overwrite caused by a faulty deployment. The table uses on-demand capacity and must remain highly available during recovery. Which feature should the data engineer enable?
Hard648Which THREE factors should a data engineer consider when choosing between Amazon RDS and Amazon DynamoDB for a new application? (Choose three.)
Hard649A data engineer needs to set up a new Amazon RDS for MySQL database for a web application. The application experiences variable read traffic and requires low read latency. The engineer needs to minimize downtime during maintenance and provide read scalability. Which configuration meets these requirements?
Easy650A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources and needs to be partitioned by year, month, day, and event type for efficient querying with Amazon Athena. Which S3 key prefix structure is most appropriate?
Easy651A data engineer is troubleshooting an AWS Glue ETL job that fails intermittently with the error 'Rate exceeded.' The job reads from an Amazon RDS for MySQL source and writes to Amazon S3. What is the MOST likely cause of this error?
Medium652A data engineer needs to share a dataset stored in Amazon S3 with another AWS account. The bucket policy currently grants access only to the owning account. What is the simplest way to grant cross-account access?
Medium653A data engineer is configuring an S3 bucket for a data lake. The engineer runs the command shown in the exhibit. What does the output indicate about the bucket?
Easy654A data engineer is designing a data pipeline that ingests streaming data from an IoT device fleet. The data must be processed in near real-time and stored in Amazon S3 for long-term analytics. Which TWO AWS services should the engineer use together to achieve this?
Medium655A data engineer must give an AWS Lambda function temporary credentials to read objects from a specific Amazon S3 prefix. The function runs in a VPC and accesses S3 through a gateway VPC endpoint. The security team forbids long-lived access keys on the function. What should the data engineer configure?
Medium656A data engineer is troubleshooting an AWS Glue ETL job that fails with an 'Access Denied' error when trying to write to an S3 bucket. The IAM role used by the job has the policy shown in the exhibit. The bucket 'my-bucket' uses S3 default encryption with AWS KMS. What is the most likely missing permission?
Hard657A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. The engineer notices that when a Glue job fails, the Step Functions execution also fails, but the engineer wants to retry the failed job up to three times with exponential backoff before failing the entire workflow. The engineer needs to implement this with minimal changes to the state machine. What should the engineer do?
Hard658A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. During a recent run, a Glue job failed due to a transient network issue, and the Step Functions execution stopped. The engineer needs to ensure that the workflow can automatically retry the failed Glue job up to three times before considering the step failed. Which Step Functions state configuration should the engineer implement?
Hard659A data engineer is using AWS Lake Formation to manage fine-grained access control on an Amazon S3 data lake. The engineer has registered the S3 bucket as a Lake Formation data location and created a table in the AWS Glue Data Catalog. The engineer needs to grant a data analyst permission to query only specific columns (customer_id, order_date) in the sales table using Amazon Athena, while hiding other columns (credit_card_number, address). The analyst uses an IAM role that has no direct S3 permissions. Which action should the engineer take?
Hard660A data engineer is configuring an Amazon Redshift cluster for a reporting workload. The team needs to load data from Amazon S3 into a Redshift table and wants the fastest possible load while keeping the data compressed. Which approach should the engineer use?
Easy661A company uses Amazon Kinesis Data Streams to ingest IoT sensor data. The data is processed by an AWS Lambda function that transforms the records and writes to an Amazon S3 bucket. Recently, the Lambda function has been failing with 'Rate exceeded' errors for the S3 PUT API calls. The data volume is 10 MB/s with average record size 2 KB. What should be done to resolve this issue?
Hard662A company stores customer transaction data in an Amazon DynamoDB table. The table has a partition key of CustomerID and a sort key of TransactionDate. The data engineering team needs to retrieve all transactions for a specific customer within a date range, and the queries must be efficient. Which DynamoDB operation should the team use?
Medium663A data engineer notices that an Amazon Kinesis Data Firehose delivery stream is failing to deliver data to an Amazon S3 bucket. The CloudWatch metrics show 'DeliveryToS3.Success' is 0 and 'S3.BucketExists' is 1. What is the MOST likely cause?
Easy664A company uses AWS Glue to run ETL jobs that transform data from Amazon S3 (Parquet) into a denormalized format for Amazon Redshift. The Glue job uses the DynamicFrame API. The job is failing with a 'MemoryError' when performing a join operation. The data is skewed on the join key. Which THREE actions can reduce memory usage and improve job stability? (Choose THREE.)
Hard665A data engineer is using Amazon Athena to query data stored in Amazon S3 in Parquet format. The engineer notices that a specific query is scanning much more data than expected, resulting in high costs and slow performance. The query filters on a column named 'event_date' which is a string in 'YYYY-MM-DD' format. The table is partitioned by 'year', 'month', and 'day' as separate string columns. The engineer wants to reduce the amount of data scanned. Which action should the engineer take?
Hard666A data engineer needs to transform JSON data from Amazon S3 into Parquet format using AWS Glue. The source files are in a bucket with thousands of small files. What is the best practice to optimize the Glue job performance?
Easy667A data engineer needs to run a daily AWS Glue ETL job that transforms data in Amazon S3. The job must start at 2:00 AM UTC every day. The engineer wants to minimize operational overhead and ensure the job runs reliably. Which approach should the engineer use?
Easy668A data engineer manages an Amazon S3 data lake with a bucket that has S3 Versioning enabled. A downstream analytics job accidentally overwrites thousands of current objects with corrupted data. The engineer must restore the previous good versions quickly and prevent the corrupted versions from being served. Which action should the engineer take?
Hard669A company needs to ingest data from an on-premises database to Amazon S3 with minimal impact on the source database. The data volume is several TB. Which AWS service is best suited for this task?
Easy670A data engineer is designing an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structures and write the output to Amazon Redshift. The engineer needs to ensure the job can handle schema evolution and efficiently process only new data on subsequent runs. (Choose two.)
Medium671A healthcare company processes patient records in near-real-time using Amazon Kinesis Data Streams. Each record contains sensitive personal health information (PHI). The data must be encrypted at rest and in transit. The company also needs to audit access to the data. The data engineer is designing the ingestion pipeline. Which combination of services and configurations meets these requirements?
Hard672A data engineer is configuring an AWS Glue ETL job to read data from an Amazon S3 bucket that contains nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer wants to optimize the job for performance and cost. Which two actions should the engineer take? (Choose two.)
Hard673A data engineer needs to ingest on-premises CSV files into Amazon S3 every hour. The files are less than 1 GB each. Which service is the most cost-effective and requires the least operational overhead?
Easy674A data engineer needs to store JSON documents that are accessed by a key-value pattern. The workload requires single-digit millisecond latency at any scale. Which AWS service is most appropriate?
Medium675A team uses Amazon Kinesis Data Analytics to process streaming data. They notice that the application's output is delayed. Which AWS service can be used to monitor the application's performance and identify bottlenecks?
Easy676A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data is in JSON format. The delivery stream is configured with a buffer size of 5 MB and a buffer interval of 60 seconds. However, the data engineer notices that S3 objects are being created with sizes much smaller than 5 MB. What is a likely cause?
Medium677Refer to the exhibit. An IAM policy for an AWS Lambda function. The Lambda function is triggered by an S3 event (object created) and needs to read from a Kinesis stream. However, the function fails with access denied when trying to read from Kinesis. What is the most likely cause?
Medium678A data engineer needs to audit data access events in Amazon S3. Which AWS service should be used to record and monitor API calls for S3 buckets?
Easy679Match each AWS storage class to its description.
Medium680A data engineer is designing a data lake on Amazon S3 that will be accessed by multiple AWS Glue ETL jobs. The engineer needs to ensure that the data is organized efficiently for querying and that sensitive columns are masked for certain users. Which TWO actions should the engineer take? (Choose TWO.)
Medium681A data engineer is setting up an Amazon Kinesis Data Firehose delivery stream to load data into Amazon Redshift. The data is coming from an application that produces JSON records. The engineer needs to transform the data to match the Redshift table schema. Which approach is the MOST cost-effective and requires the least operational overhead?
Easy682A company is building a data lake on Amazon S3 and wants to ingest data from multiple AWS services (CloudTrail, VPC Flow Logs, and ALB logs). The data should be stored in a central S3 bucket with a common partitioning scheme. Which service can be used to collect and centralize this data with minimal configuration?
Medium683A data engineer manages an Amazon S3 data lake with millions of small JSON files. To improve query performance with Amazon Athena, the engineer wants to compact these files into larger Parquet files. The engineer must also minimize ongoing storage costs. Which solution should the engineer implement?
Medium684A data engineer needs to store semi-structured JSON data that is accessed infrequently but must be retrievable within minutes. The data is generated by IoT devices and each object is about 500 KB. The engineer wants the most cost-effective storage solution. Which AWS service should be used?
Medium685A data engineer is configuring an Amazon Redshift cluster to encrypt data at rest. The company policy requires that encryption keys be stored in AWS CloudHSM. Which integration should the engineer use to meet this requirement?
Medium686A data engineer is managing an Amazon S3 data lake that contains raw JSON data. The engineer needs to optimize the data lake for query performance and cost when using Amazon Athena. The data is currently stored in a single S3 prefix without partitioning, and queries often filter on `event_type` and `event_date`. The engineer wants to implement best practices for Athena. Which TWO actions should the engineer take? (Choose two.)
Medium687A data engineer needs to transfer 50 TB of historical data from an on-premises HDFS cluster to Amazon S3. The network bandwidth is limited to 100 Mbps. The transfer must be completed within one week. Which service should be used?
Easy688A company stores sensitive data in Amazon S3 and requires that all data be encrypted at rest. The data is accessed by multiple AWS services. Which solution meets the encryption requirement with the LEAST operational overhead?
Medium689A company uses AWS Glue to catalog data in Amazon S3. The data arrives in Parquet format, but the crawler fails to update the schema when new columns are added. What is the most likely cause?
Medium690A company stores application logs in Amazon S3 and uses AWS Glue crawlers to populate the AWS Glue Data Catalog. A data engineer needs to query the logs with Amazon Athena. The logs are partitioned by year/month/day in S3, but Athena queries are scanning all partitions and returning errors about missing partitions. What should the engineer do to enable partition pruning?
Easy691A media company stores millions of small JSON files in an Amazon S3 bucket and queries them with Amazon Athena. Analysts report that queries scan far more data than expected, and costs are rising. The data engineer confirms that the files are uncompressed, use no partitioning, and are stored as newline-delimited JSON. Which change will MOST reduce the data scanned per query?
Hard692A data engineer runs an AWS Glue job that reads Parquet files from Amazon S3 partitioned by year/month/day and writes to another S3 prefix. The job currently processes all historical partitions on every run, causing long runtimes and high cost. The engineer wants subsequent runs to process only new data. Which configuration should the engineer apply?
Hard693A data engineer is designing a data lake on Amazon S3 for a retail company. The company ingests point-of-sale data as small JSON files every few minutes, totaling about 5 GB per day. Analysts query the data with Amazon Athena, and costs are rising due to many small files and full scans. The engineer wants to reduce Athena query costs and improve performance while keeping the data in S3. Which TWO actions should the engineer take? (Choose two.)
Hard694A company ingests application logs into Amazon S3 through Amazon Kinesis Data Firehose. The logs arrive as newline-delimited JSON, and analysts query them with Amazon Athena. Query performance is poor because the JSON files are small and uncompressed. The engineer must improve Athena query performance while keeping the raw JSON available for a downstream legacy system. Which change should the engineer make?
Medium695A data engineer is using AWS Glue to run an ETL job that reads from an Amazon RDS for PostgreSQL database and writes to Amazon S3. The job is configured with a JDBC connection to RDS. The engineer notices that the job fails intermittently with a 'Connection timed out' error. The RDS instance is in a private subnet, and the Glue job has been configured with a VPC connection. Which action should the engineer take to resolve the timeout issue?
Hard696A company wants to enforce that all data written to an S3 bucket is encrypted with a customer-managed AWS KMS key. The data engineer has created the KMS key and attached an S3 bucket policy. However, users are still able to upload objects without specifying the KMS key. What is the most likely cause?
Easy697A company uses AWS Glue to process streaming data from Amazon Kinesis Data Streams. The job fails intermittently with a 'MemoryError'. What is the MOST likely cause?
Medium698A data engineer is using AWS Glue to process a large dataset stored in Amazon S3 in Parquet format. The Glue job performs a join between two tables and writes the result back to S3. The engineer notices that the job is running slowly and consuming excessive DPU hours. The job has 10 workers of type G.1X. Which action should the engineer take to improve performance and reduce cost?
Hard699A data engineer needs to migrate an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB and has a continuous stream of write operations. The migration should minimize downtime. Which AWS service should be used?
Medium700A company stores sensitive financial data in Amazon S3 and requires that all data be encrypted at rest using customer-managed keys. A data engineer configures the S3 bucket to use SSE-KMS with a customer-managed KMS key. The security team now wants to audit all API calls that use the KMS key to decrypt data. Which AWS service should the engineer use to capture and review these KMS API calls?
Easy701A data engineer needs to restrict access to an S3 bucket so that only users from a specific AWS account can read objects. Which S3 bucket policy element should be used?
Easy702A data engineer needs to ingest data from an Amazon DynamoDB table into an Amazon S3 data lake. The table is updated frequently and the engineer must capture all item-level changes in near real-time without impacting table performance. The ingested data must be stored in a format that preserves the change type (INSERT, MODIFY, REMOVE). Which solution meets these requirements with the LEAST operational overhead?
Medium703A company runs an Amazon RDS for PostgreSQL instance for an OLTP application. The database size is 500 GB. The company wants to minimize downtime during backups and ensure point-in-time recovery (PITR) for the last 7 days. Which TWO features should the company use? (Choose TWO.)
Hard704A company uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data is transformed using an AWS Lambda function. Some records fail transformation and are lost because the Lambda function throws an exception. The data engineer needs to capture the failed records for analysis without affecting the pipeline. What should the engineer do?
Hard705A company uses Amazon Redshift for data warehousing. They notice that query performance has degraded over time. Which maintenance operation should be performed to improve performance?
Easy706A company needs to transfer 20 TB of historical data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited and the transfer must complete within one week. Which service should the company use?
Easy707A company stores sensitive data in an Amazon S3 bucket. A compliance requirement mandates that all data must be encrypted at rest with a key that is automatically rotated every year. The company also needs to maintain an audit trail of who used the key. Which solution meets these requirements?
Medium708A company uses AWS Glue to transform data in S3. The transformation job reads Parquet files, filters rows, and writes to another S3 bucket. The job takes longer than expected. Which change would MOST likely reduce the job execution time?
Medium709A company uses AWS Glue to process data from Amazon RDS MySQL into Amazon S3. The Glue job uses a JDBC connection and runs on a schedule. Recently, the job has been failing with a 'Communications link failure' error. The RDS instance is in a private subnet. Which troubleshooting step should the data engineer take FIRST?
Hard710A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream and writes results to an Amazon S3 bucket. Recently, the application has been failing with 'ResourceNotFoundException' for the S3 bucket. What is the MOST likely cause?
Medium711A data engineer is designing a solution to ingest streaming data from Amazon Kinesis Data Streams into an Amazon Redshift cluster for near-real-time analytics. The engineer needs to ensure that data is loaded efficiently and that the Redshift cluster can handle the ingestion load without impacting query performance. Which approach should the engineer use?
Medium712A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job needs to run daily and process only new data since the last run. The data in S3 is partitioned by date in the format `year=YYYY/month=MM/day=DD/`. Which feature should the data engineer use to track processed partitions and avoid reprocessing old data?
Medium713Refer to the exhibit. A data engineer runs this CLI command on an S3 bucket. The data is ingested from multiple sources. Which AWS service would be best to process these files in a single batch transformation?
Easy714A data engineer must ingest data from an Amazon DynamoDB table into an Amazon S3 bucket for analytics. The table receives continuous writes, and the ingestion must capture changes with minimal latency. The engineer wants a fully managed solution that requires no server provisioning. Which approach should the engineer use?
Medium715A company wants to audit API calls made to its Amazon S3 buckets. Which AWS services can be used to achieve this? (Choose TWO.)
Easy716A data engineering team is designing a batch processing workflow using AWS Glue. The job reads from an S3 bucket, transforms data, and writes to another S3 bucket. The job runs daily and processes new data incrementally. Which THREE features should they use to optimize performance and cost?
Hard717A company is designing a data ingestion pipeline for real-time IoT sensor data. The data volume peaks at 10,000 messages per second. The pipeline must process messages in order per sensor and persist raw data to Amazon S3 for archival. Which TWO services should be used together to meet these requirements? (Choose TWO.)
Easy718A company stores sensitive data in Amazon S3 and uses AWS KMS customer-managed keys for encryption. The security team wants to monitor and audit all KMS API calls that involve the key, including who used the key and when. They also want to receive alerts if the key is used by an unauthorized principal. Which AWS service should the data engineer use to meet these requirements?
Medium719A data engineer is designing a pipeline to ingest change data capture (CDC) events from an Amazon RDS for PostgreSQL database into Amazon S3. The CDC events are captured using AWS DMS. The data must be available for querying within 5 minutes of the change. Which approach meets these requirements?
Medium720A company uses AWS Lake Formation to manage access to data in a data lake. A new data engineer has been granted SELECT permission on a table but receives an 'AccessDeniedException' when querying via Amazon Athena. The table is registered in Lake Formation and the data is encrypted with SSE-KMS. Which of the following is the MOST likely cause?
Hard721A company uses AWS Glue DataBrew to clean and normalize data from an Amazon S3 bucket before loading it into Amazon Redshift. The data contains PII such as social security numbers. A compliance policy requires that PII be masked in the DataBrew output. Which DataBrew transformation should the engineer use to replace the last four digits of each SSN with asterisks?
Medium722An IAM policy includes the above resource ARN for CloudWatch Logs. A data engineer needs to allow a Lambda function to create log streams and put logs to the log group 'my-log-group'. However, the Lambda function is failing with access denied. What is the issue?
Easy723A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it with Amazon Athena. The engineer wants to minimize query costs and improve performance. Which action should be taken?
Easy724A data engineer is troubleshooting slow query performance on an Amazon Redshift cluster. The cluster has 10 nodes and is using automatic distribution style. The engineer suspects that data distribution is causing excessive data movement. Which steps should the engineer take to diagnose and resolve the issue? (Choose THREE.)
Hard725A company uses Amazon S3 to store log files from multiple applications. The logs are written in JSON format. A data engineer wants to use Amazon Athena to query these logs. The logs are stored in a bucket with the following structure: 's3://logs/app1/date=2021-01-01/'. The engineer creates an Athena table with partitions. However, when querying, Athena returns zero results for partitions that exist. The engineer has run MSCK REPAIR TABLE to add partitions. What is the most likely cause of the issue?
Easy726A data engineer is building a governed data lake in AWS Lake Formation. The security team wants to detect sensitive data such as credit card numbers in newly registered S3 tables and automatically apply column-level access restrictions to those columns. Which TWO actions should the engineer take to meet these requirements? (Choose two.)
Hard727An application uses the 'orders' DynamoDB table with the schema and provisioned throughput shown in the exhibit. The application frequently queries by customer_id (range key) without specifying the order_id (partition key). What is the most likely impact on performance?
Hard728A company is using Amazon RDS for MySQL with Multi-AZ deployment. The database size is 2 TB and the workload is read-heavy. To improve read performance, which option should be used?
Medium729A data engineer needs to grant an IAM user read-only access to an S3 bucket named 'data-lake'. Which IAM policy statement should be used?
Easy730A data engineer is using AWS Glue to transform data from an Amazon S3 bucket. The Glue job reads JSON files, applies complex transformations, and writes the output to another S3 bucket in Parquet format. The job runs daily and must complete within a 2-hour window. The engineer notices that the job is taking longer than expected and wants to optimize performance. The source data is partitioned by date, and the job uses a dynamic frame. Which optimization should the engineer implement to improve performance?
Hard731A data pipeline uses AWS Glue to process data from Amazon S3 and write results to Amazon Redshift. The pipeline fails intermittently with the error 'S3ServiceException: Access Denied'. The IAM role used by Glue has permissions to read from the S3 bucket. What is the most likely cause of this error?
Medium732Refer to the exhibit. A data engineer is troubleshooting an AWS Glue ETL job that fails with an access denied error when writing to S3. The IAM role attached to the Glue job has the policy shown. What is the most likely cause of the error?
Medium733A data engineer is using AWS Glue to catalog data stored in Amazon S3. The engineer needs to run an AWS Glue ETL job that reads from a large dataset in Parquet format, performs transformations, and writes the output to Amazon Redshift. The job must handle data skew and optimize performance. Which AWS Glue feature should the engineer use to address data skew during the join operation?
Hard734A data engineer is troubleshooting a slow-running query on Amazon Redshift. The query scans a large table but returns few rows. Which diagnostic step should be taken first?
Medium735A company is migrating a large Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime and preserve data consistency. Which THREE AWS services or features should be used?
Hard736A data engineer manages an AWS Glue ETL job that reads JSON files from Amazon S3, transforms the data, and writes to an Amazon Redshift table. The job recently started failing with the error 'Communication link failure: connection reset'. The Redshift cluster is healthy and the IAM role used by the Glue job has the necessary permissions. The engineer notices that the job runs longer than before and the Redshift cluster's WLM queue is often full. Which action should the engineer take to resolve the failure?
Medium737A data engineer must transform data in an AWS Glue ETL job. The transform requires calling an external REST API for each record to enrich the data. The Glue job runs on AWS Glue 4.0 with Python. The engineer wants to minimize the number of API calls and improve performance. Which approach should the engineer take?
Medium738A company stores sensitive data in Amazon Redshift. The security team requires that all data in the cluster be encrypted at rest using a customer managed key in AWS KMS, and that the key be rotated annually. The data engineer needs to configure the Redshift cluster accordingly. Which action should the engineer take?
Medium739A data engineer is using Amazon Redshift to store sales data. The engineer needs to ensure that the data is encrypted at rest and that encryption keys are managed by AWS. The engineer also wants to minimize administrative overhead. Which Redshift encryption option should the engineer use?
Easy740A company stores sensitive data in Amazon S3 and needs to ensure that data is encrypted at rest. Which AWS service can be used to manage the encryption keys?
Easy741A company is using Amazon RDS for PostgreSQL with Multi-AZ deployment. The primary instance fails and a failover occurs. After the failover, the application cannot connect to the database. What is the MOST likely cause?
Medium742A data engineer is designing a data ingestion pipeline to load clickstream data from an Amazon S3 bucket into an Amazon Redshift cluster. The data arrives in 5-minute batches. Which TWO actions should the engineer take to ensure data consistency and avoid duplicates? (Select TWO.)
Medium743A company uses AWS Glue to process sensitive customer data stored in S3. The security team requires that all data be encrypted at rest using a customer-managed KMS key and that access to the key be auditable. Which solution meets these requirements?
Medium744A company has a data lake in Amazon S3 with millions of objects. The security team wants to enforce that all objects are encrypted with a specific customer-managed KMS key. The data engineer configures an S3 bucket policy to deny PutObject if the encryption is not set to that key. However, some existing objects are not encrypted with that key. What is the most efficient way to remediate the existing objects?
Hard745A company is storing sensitive user data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using a customer-managed key stored in AWS KMS. The bucket policy must deny any PUT request that does not include the appropriate encryption header. Which bucket policy condition key should be used?
Medium746A data engineer is managing an Amazon Kinesis Data Firehose delivery stream that writes to an Amazon S3 bucket. The engineer notices that some records are being delivered to the S3 bucket with a prefix of 'errors/' instead of the intended 'data/' prefix. The Firehose stream is configured with an AWS Lambda function for data transformation, and the S3 bucket has a lifecycle policy that transitions objects to Glacier after 30 days. Which TWO actions should the engineer take to ensure that only successfully transformed records are delivered to the 'data/' prefix and that failed records are handled appropriately? (Choose two.)
Medium747A data engineer is troubleshooting a data pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The engineer notices that the S3 bucket contains many small files (less than 1 MB). This is causing performance issues in downstream processing. What is the BEST way to reduce the number of small files?
Medium748A company is designing a data lake on Amazon S3 for analytics. The data includes sensitive personally identifiable information (PII). Which TWO actions should the company take to protect the data? (Choose TWO.)
Medium749A data engineer runs the above CLI command to describe the DynamoDB table 'Orders'. The table has a partition key 'OrderID' and sort key 'CustomerID'. Which query operation is most efficient for retrieving all orders for a specific customer?
Medium750A company uses AWS Lake Formation to manage data lake permissions. The data engineer notices that a user with SELECT permission on a table can also query the underlying data in Amazon S3 directly. How can the engineer enforce that access to the S3 data is only through Lake Formation?
Hard751A company is ingesting log files from multiple EC2 instances into Amazon S3 using the CloudWatch agent. The logs are delivered to a CloudWatch Logs group, and a subscription filter sends them to a Lambda function for transformation, then to Firehose. The Firehose stream is configured with a buffer interval of 60 seconds and buffer size of 5 MB. The logs are critical and must be available in S3 within 5 minutes. What is the most cost-effective way to reduce the delivery latency?
Medium752A data engineering team notices that an AWS Glue ETL job, which processes hourly data from an S3 bucket, is taking progressively longer to run. The job reads Parquet files partitioned by date and hour. Which action is MOST likely to improve the job's performance?
Medium753A company runs a nightly batch ETL job using AWS Glue to transform data from Amazon RDS for MySQL to Amazon S3. The job reads 100 tables and writes Parquet files partitioned by date. Recently, the job started failing with 'ThrottlingException' from the RDS database. The data volume has increased, and the Glue job is reading large tables without any filtering. The job uses a single Glue job with multiple Spark executors. The engineer needs to reduce the load on the RDS database while maintaining the same processing time. What should the engineer do?
Hard754Refer to the exhibit. A data engineer applies this bucket policy to an S3 bucket. A user within the 10.0.0.0/24 IP range attempts to upload an object to the bucket using an HTTP (non-HTTPS) request. What is the outcome?
Hard755A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 bucket containing nested JSON, flattens arrays using Relationalize, and writes Parquet to another S3 bucket. The job must run only when new objects land in the source bucket. The engineer wants to avoid unnecessary job runs and minimize cost. Which approach should the engineer use?
Hard756A data engineer maintains an Amazon Redshift cluster where a nightly COPY job loads data into a large fact table. After the load, analysts run queries that filter on a `sale_date` column and join to a small dimension table. Query performance degrades over time as the fact table grows. The engineer wants to improve performance for these recurring queries without changing the query text. Which combination of actions should the engineer take?
Hard757A data engineer is using AWS Glue DataBrew to prepare a dataset stored in Amazon S3. The dataset contains missing values, inconsistent date formats, and duplicate rows. The engineer needs to clean the data and produce a transformed output for downstream analytics. (Choose two.)
Medium758A company runs a data ingestion pipeline that uses AWS Glue to read 500 GB of JSON files from an S3 bucket (s3://raw-data/) every hour. The Glue ETL job transforms the data and writes Parquet files to another S3 bucket (s3://processed-data/). The job is triggered by a time-based CloudWatch Events rule. Recently, the job has started taking over 2 hours to complete, causing delays in downstream processes. The data volume has been consistent, and no changes have been made to the job code or infrastructure. The S3 bucket 's3://raw-data/' receives new files continuously, but the Glue job reads all files in the bucket each run (no incremental processing). The engineer suspects that the job is reprocessing old data. Which action should the engineer take FIRST to reduce the job duration?
Hard759A data engineer is designing a data ingestion pipeline to load data from an on-premises Oracle database to Amazon S3. The pipeline should capture changes in near real-time (within minutes) and minimize impact on the source database. The source table has a 'last_modified' timestamp column. Which service combination would meet these requirements?
Medium760A data engineer is designing a data ingestion pipeline for clickstream data from a mobile app. The data volume varies, with occasional spikes up to 10 MB/s. The pipeline must persist the raw data in Amazon S3 and make it available for near-real-time analytics via Amazon Athena. Which combination of services minimizes cost and operational overhead?
Hard761A data engineer manages an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to an Amazon Redshift table. The job runs daily and recently started failing with the error 'Unable to find a suitable security group for the connection'. The Glue connection is configured with a VPC, subnet, and security group. The engineer verifies that the IAM role has the necessary permissions and the S3 bucket is accessible. What is the most likely cause of this error?
Medium762A company is using Amazon EMR to process data stored in Amazon S3. The S3 bucket is configured with a bucket policy that denies access unless the request includes a specific tag. The EMR cluster's IAM role has s3:GetObject permission. However, the EMR job fails to read data from S3. What is the most likely cause?
Hard763A data engineer maintains an Amazon DynamoDB table that stores IoT telemetry. The table uses a partition key of deviceId and a sort key of timestamp, with on-demand capacity. A few very active devices generate millions of writes per hour while thousands of other devices write sporadically. The engineer observes throttling on writes for the active devices and wants to reduce it with the least application change. Which action should the engineer take?
Hard764A company is using Amazon EMR with Kerberos authentication. They want to ensure that data in transit between EMR cluster nodes is encrypted. Which configuration should be applied?
Hard765A data engineer is migrating an on-premises MongoDB database to Amazon DocumentDB. Which migration strategy minimizes downtime?
Medium766A data engineer needs to ingest streaming data from thousands of IoT devices and immediately process each record with minimal latency. Which AWS service should be used as the ingestion point?
Easy767A company uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data is transformed using an AWS Lambda function. Recently, the transformation errors have increased due to Lambda timeouts. The data engineer needs to diagnose and resolve the issue without losing data. What should the engineer do?
Hard768A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Apache Parquet format and partitioned by year, month, and day. The engineer notices that AWS Glue crawlers are taking a long time to run and are not correctly identifying new partitions. The engineer needs to improve the crawler performance and ensure new partitions are added automatically. Which action should the engineer take?
Hard769A data engineering team is building an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another. The security team mandates that both read and write operations use a customer-managed AWS KMS key so they can audit key usage. Which configuration should the data engineer apply to the Glue job to meet this requirement?
Medium770A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing sensitive customer records. The job must remove all columns flagged as PII before writing to a target S3 location. The PII columns are not known in advance and vary by file. Which approach should the engineer use to ensure the PII is removed dynamically?
Medium771A company needs to ingest data from a MySQL database into Amazon S3 in near real-time. The database is running on EC2. The data engineer wants to minimize the impact on the source database. Which service should be used?
Easy772A financial services company stores transaction data in Amazon RDS for PostgreSQL. The company requires that all changes to the database be logged for audit purposes, including before and after images of updated rows. Which feature should the data engineer enable?
Hard773A data engineer needs to audit all AWS KMS key usage in the account. Which AWS service should be used to record KMS API calls?
Easy774A data engineer needs to store archival data that is rarely accessed but must be retained for 7 years. The data should be retrievable within 12 hours. Which Amazon S3 storage class is MOST cost-effective?
Easy775A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data should be stored cost-effectively. Which storage solution meets these requirements?
Easy776A data engineer is monitoring an Amazon Kinesis Data Analytics for Apache Flink application that processes streaming data. The application is falling behind (increasing 'MillisBehindLatest') and the CPU utilization of the Flink task managers is consistently above 80%. Which THREE actions should the engineer take to improve performance? (Choose THREE.)
Medium777A data engineer is setting up a new Amazon Redshift cluster for a data warehouse. The engineer wants to ensure data durability and high availability. Which THREE features should the engineer consider? (Choose three.)
Easy778A company stores sensitive data in Amazon S3. They need to ensure that all objects are encrypted at rest. Which approach meets this requirement with minimal effort?
Medium779A data engineer needs to create a table in Amazon Athena that reads JSON data stored in Amazon S3. The JSON records are stored in a single file, one JSON object per line. The engineer wants Athena to automatically discover the schema and create the table without manually defining columns. Which AWS service or feature should the engineer use?
Medium780A data engineer is designing a data lake on Amazon S3. The data includes customer PII that must be encrypted at rest. The company also requires that the encryption keys be rotated automatically every year. Which encryption solution should the engineer use?
Easy781A data engineer is designing a data pipeline that processes streaming data. The pipeline must be able to handle duplicate records and ensure exactly-once processing semantics. Which THREE AWS services or features should the engineer consider? (Choose three.)
Easy782A data engineer is troubleshooting an AWS Glue ETL job that fails with a 'java.lang.OutOfMemoryError: Java heap space' error. The job processes a 50 GB Parquet file from an S3 bucket. The job uses a G.1X DPU (16 GB memory) and default parameters. Which action should the engineer take to resolve the issue?
Medium783A company uses Amazon DynamoDB with provisioned capacity. During a sales event, write traffic spikes and some requests receive ProvisionedThroughputExceeded exceptions. The reads are within limits. The data engineer needs to minimize latency for the spike without manual intervention. Which solution is MOST cost-effective?
Hard784A data engineer is designing a streaming pipeline that ingests IoT sensor data from 10,000 devices. Each device sends a 1 KB message every second. The data must be processed in near real-time and stored in S3 for analytics. Which combination of services provides the most cost-effective solution?
Hard785A company wants to audit all changes to IAM policies in their AWS account. Which combination of services should be used to achieve this?
Hard786A data engineer creates an Amazon DynamoDB table using the CloudFormation snippet in the exhibit. The application writes 200 items per second to the table. The engineer notices that many write requests are being throttled. What is the MOST likely reason?
Easy787A company stores critical financial data in Amazon DynamoDB. To meet compliance requirements, the data must be encrypted at rest with a customer-managed key. Which solution should the data engineer implement?
Easy788A data engineer is designing a data lake on Amazon S3 using AWS Lake Formation. The engineer needs to grant fine-grained access to specific columns and rows of a table to different analysts. Which two actions should the engineer take to meet these requirements? (Choose two.)
Hard789A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes only new data added since the last run. The engineer notices that the job is reprocessing old data, increasing runtime and cost. Which AWS Glue feature should be enabled to ensure only new data is processed?
Medium790A data engineer is using Amazon Redshift and needs to improve the performance of complex queries that join large tables. The engineer has already set the distribution style to KEY on the join columns. What additional step should the engineer take to optimize the join performance?
Medium791A company is using Amazon DynamoDB with on-demand capacity for a gaming application. During a new game launch, write traffic spikes to 50,000 writes per second, but the application experiences throttling. The DynamoDB table has a partition key of 'game_id' and a sort key of 'timestamp'. What is the MOST likely cause of throttling?
Hard792A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for ad-hoc querying using SQL. Which AWS service should be used?
Medium793A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job must run incrementally and process only new data since the last run. The source data is partitioned by date in S3, and new partitions are added daily. Which AWS Glue feature should the engineer enable to track previously processed data?
Medium794A company is using Amazon S3 as a data lake. The data engineer needs to ensure that all objects uploaded to a specific bucket are automatically replicated to a bucket in another AWS Region for disaster recovery. Which configuration should the engineer implement?
Easy795A company needs to enforce encryption in transit for all data moving between its Amazon S3 bucket and a fleet of Amazon EC2 instances. The data is accessed via S3 API calls over the internet. Which configuration ensures encryption in transit?
Medium796A company is using Amazon DynamoDB with auto scaling enabled. During a marketing campaign, write traffic spikes, and some write requests fail with ProvisionedThroughputExceededException. The auto scaling policy has a target utilization of 70% and a maximum capacity that is high enough. What is the most likely cause of the throttling?
Hard797A company is building a data pipeline that ingests sensitive customer data from an on-premises database into Amazon S3 using AWS DMS. The data must be encrypted at rest in S3 and in transit. The security team requires that the encryption keys be managed by the company (not AWS). Which TWO actions should the data engineer take to meet these requirements? (Choose TWO.)
Medium798A company is designing a data lake on Amazon S3. The data includes personal identifiable information (PII). The data engineer must ensure that only authorized users can access the data, and that access is logged for auditing. Which combination of services should the data engineer use?
Hard799A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains a column with inconsistent date formats (for example, '2023-01-15', '01/15/2023', and '15-Jan-2023'). The engineer needs to standardize all values to ISO 8601 format and then write the cleaned data back to S3. Which approach should the engineer use?
Hard800A data engineer is troubleshooting an access denied error when an AWS Lambda function tries to decrypt an object encrypted with the KMS key 'abc123'. The Lambda function's execution role has the above policy attached. What is the likely cause of the error?
Hard801A company uses Amazon DynamoDB as the primary data store for a gaming application. The application stores user profiles and game state. During peak hours, the application experiences throttling on writes to the UserProfiles table. The table's read capacity is underutilized. Which solution should resolve the write throttling?
Easy802A data engineer is designing a data lake on Amazon S3. The data consists of sensitive personally identifiable information (PII) that must be encrypted at rest. The company requires that encryption keys be rotated every 90 days and that access to the keys be logged. Which encryption solution meets these requirements?
Medium803A financial services company uses a multi-account AWS Organization with hundreds of accounts. The data engineering team needs to enable cross-account access to an encrypted S3 bucket in the data lake account (account ID 111111111111) for a Glue ETL job running in the analytics account (account ID 222222222222). The S3 bucket uses AWS KMS customer managed key (CMK) for server-side encryption (SSE-KMS). The Glue job fails with an AccessDenied error when trying to read data from the bucket. The IAM roles in both accounts have the necessary S3 permissions and the bucket policy allows access from the analytics account. What is the most likely cause of the failure?
Hard804A data engineer is troubleshooting a Lambda function that reads from a Kinesis Data Stream, processes records, and writes to a Kinesis Data Firehose delivery stream. The Firehose delivery stream is configured to deliver data to an S3 bucket. The Lambda function is failing with an access denied error. The IAM policy attached to the Lambda execution role is shown in the exhibit. Which permission is missing?
Hard805A data engineer is building a data lake on Amazon S3 and must choose the optimal file format for a dataset that is queried by Amazon Athena. The queries typically select a few columns from wide tables containing hundreds of columns, and the data volume is in terabytes. The engineer wants to minimize query scan costs and improve performance. Which file format should the engineer use?
Medium806A company wants to ingest streaming data from thousands of IoT devices into AWS for real-time processing. Each device sends JSON payloads of about 2 KB at a rate of 1 message per second. The data must be processed with a durable, ordered stream per device. Which service should the company use as the ingestion layer?
Easy807Refer to the exhibit. A data engineer is configuring an AWS Lambda function to process records from a Kinesis stream. The function is set up with an event source mapping, but no records are being processed. The Lambda function's IAM role has the policy shown. What is the most likely reason for the issue?
Hard808A CloudFormation template includes this IAM policy for a cross-account S3 upload use case. What is the purpose of the condition?
Hard809A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data is then processed by an AWS Lambda function that transforms the records and writes them to an Amazon S3 bucket. Recently, the Lambda function has been timing out and the S3 bucket is not receiving all expected data. The Kinesis stream is not throttling and has sufficient shards. Which step should the company take to resolve this issue?
Medium810A data engineer is troubleshooting a failed AWS Glue ETL job that reads from an S3 bucket. The job logs show the following error: 'java.lang.RuntimeException: java.lang.ClassNotFoundException: Class org.apache.hadoop.fs.s3a.S3AFileSystem not found'. Which TWO actions will resolve this issue?
Hard811Refer to the exhibit. An IAM policy includes the above statement to allow decryption of a KMS key under specific conditions. What does this policy allow?
Easy812A CloudFormation template defines an AWS Glue job. The job fails during execution with the error 'Unable to locate script: s3://scripts-bucket/etl-script.py'. The S3 bucket 'scripts-bucket' exists and the script file is present. What is the most likely cause?
Hard813A data engineer runs the above command and gets the output. What does the 'MFADelete' setting imply?
Medium814A data engineer needs to ensure that all data in an S3 bucket is encrypted at rest. The bucket currently contains unencrypted objects from past uploads. Which action will encrypt these existing objects without re-uploading them?
Easy815A company has an AWS Glue ETL job that reads data from an S3 bucket, transforms it, and writes to another S3 bucket. The security team requires that data in transit between the Glue job and S3 be encrypted using TLS. The Glue job runs in a VPC with a VPC endpoint for S3. Which configuration ensures TLS encryption for all data transfer?
Hard816A company uses AWS Data Pipeline to copy data from DynamoDB to S3 daily. Recently, the pipeline started failing with 'ThrottlingException' errors. The DynamoDB table has on-demand capacity. Which action should be taken to resolve the issue?
Medium817A company runs an Amazon RDS for PostgreSQL instance that stores financial data. The company requires point-in-time recovery (PITR) with a retention period of 35 days. Additionally, the company needs to create a new database from a specific snapshot every night for testing. Which combination of actions should the data engineer take to meet these requirements?
Hard818A data engineer needs to store log files from multiple applications in a centralized location. The logs are generated in JSON format and each log entry is about 1 KB. The engineer needs to query the logs occasionally using SQL-like queries. Which AWS service is most appropriate?
Easy819A company wants to centrally manage access to multiple AWS accounts for its data engineers. The company already uses AWS Organizations. Which AWS service should be used to define fine-grained permissions across accounts?
Easy820A company uses Amazon Redshift for data warehousing. The security team requires that all data in transit between the Redshift cluster and clients be encrypted. Which feature should be enabled?
Easy821A company stores sensitive data in Amazon S3 and uses AWS Lake Formation to manage fine-grained access control. A data engineer notices that users are able to access data in S3 directly via the AWS Management Console, bypassing Lake Formation permissions. What should the engineer do to enforce Lake Formation access controls for all access methods?
Easy822A company is using Amazon EMR to process large datasets stored in Amazon S3. The data engineer wants to reduce the time it takes to read data from S3 by optimizing the data format. Which file format should the engineer recommend?
Easy823A data engineer is designing a data ingestion pipeline for real-time financial transactions. The pipeline must ensure exactly-once processing semantics and must handle duplicate records that may occur due to retries. Which combination of AWS services can achieve exactly-once processing?
Hard824Which TWO are benefits of using Amazon S3 Object Lock? (Choose TWO.)
Hard825A company uses AWS Lake Formation to manage data lakes on Amazon S3. The data engineer needs to grant a data analyst access to query specific columns in a table using Amazon Athena, but deny access to columns containing personally identifiable information (PII). Which Lake Formation feature should be used?
Hard826A company runs a data lake on Amazon S3 with AWS Glue and Amazon Athena. The data engineer notices that queries are slow and scanning large amounts of data. Which THREE actions should the engineer take to optimize query performance and reduce costs?
Hard827A company runs a data pipeline that uses AWS Lambda to process files uploaded to an S3 bucket. Recently, some files have been processed multiple times. The Lambda function is triggered by S3 event notifications. What is the MOST likely cause of duplicate processing?
Easy828A company needs to ingest data from multiple SaaS sources (e.g., Salesforce, Marketo) into Amazon S3 for analytics. Which AWS service is designed for this purpose?
Easy829A data engineer needs to ensure that an Amazon Redshift cluster only accepts encrypted connections. Which parameter should be modified?
Easy830A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak hours. The cluster uses a single node type and has no concurrency scaling enabled. Analysis shows that many long-running queries are queued behind short ad-hoc queries, causing delays for critical reports. The engineer needs to ensure that critical reports run promptly without affecting ad-hoc queries. Which solution meets these requirements?
Hard831A data engineer needs to set up a disaster recovery solution for an Amazon RDS for MySQL database. The database must be available in another AWS Region with minimal data loss. What is the simplest approach?
Easy832A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a table. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition and ensure that the load is efficient and cost-effective. Which method should the engineer use?
Hard833A company is building a data lake on Amazon S3 and needs to ingest data from various on-premises sources. Which TWO AWS services can be used to transfer data securely over the internet?
Easy834A company is building a data lake on Amazon S3. They need to ingest data from multiple sources, including relational databases, streaming data, and log files. Which THREE AWS services can be used to ingest data into the data lake?
Medium835A data engineer is setting up an Amazon RDS for MySQL database. The compliance team requires that all data at rest be encrypted. What must the engineer do to enable encryption for this database?
Easy836A data engineer needs to set up a new Amazon RDS for PostgreSQL database for a production workload. The database must be highly available and resilient to a single Availability Zone failure. Which configuration should the engineer choose?
Hard837A data engineer must load a 50 GB uncompressed CSV file from Amazon S3 into an Amazon Redshift cluster using the COPY command. The load is taking a long time and the engineer wants to improve performance. Which action should the engineer take?
Easy838A data engineer needs to transfer 10 TB of data from an on-premises data center to Amazon S3. The network bandwidth is limited to 100 Mbps, and the data transfer must be completed within 5 days. What is the most cost-effective solution?
Easy839A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The engineer notices that some records are being delivered to S3 with a delay of several minutes, and sometimes records are missing. The Firehose stream is configured with a buffer size of 5 MB and a buffer interval of 300 seconds. The engineer wants to reduce latency and ensure all records are delivered. Which action should the engineer take?
Medium840A company is using Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key. The data engineer must ensure that only a specific IAM role can decrypt the data. Which policy should the data engineer attach to the KMS key?
Hard841A data engineer is troubleshooting a slow Amazon Redshift query that joins a large fact table with several dimension tables. The EXPLAIN plan shows a hash join on the distribution key, but the query still runs slowly. The fact table is distributed by KEY(column_x) and the dimension tables are distributed ALL. The engineer notices that the fact table has a high number of rows with the same value in column_x. What is the most likely cause of the slow performance?
Hard842A company uses AWS Glue to run ETL jobs that process data from Amazon RDS to Amazon S3. The jobs run nightly and take 3 hours to complete. The data volume is growing by 20% each month. The engineer needs to reduce job runtime and cost. The source RDS is a db.r5.large instance. Which approach would be MOST effective?
Hard843Refer to the exhibit. An IAM policy is attached to a user who needs to read objects from the 'example-bucket' S3 bucket. The user reports being unable to read any object under the 'confidential/' prefix. What is the reason for this access issue?
Medium844A data engineer manages an AWS Glue job that processes JSON files from Amazon S3 and writes to Amazon Redshift. The job fails with the error "Unable to find a suitable JDBC driver". The engineer has verified that the Glue connection to Redshift is configured correctly and the IAM role has the necessary permissions. What is the most likely cause of this error?
Medium845A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data must be transformed and stored in Amazon S3 for batch analytics. The engineer wants to use AWS Lambda for transformation. Which TWO configurations are required? (Choose two.)
Medium846A data engineer is configuring AWS Glue to crawl a dataset stored in Amazon S3 and populate the AWS Glue Data Catalog. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS. The engineer has already configured the Glue crawler to use a connection with the appropriate VPC settings. What additional step must the engineer take to enforce encryption in transit?
Hard847A data engineer is configuring an Amazon S3 bucket that will receive raw clickstream files from a mobile application. The engineer must ensure that the objects are protected against accidental overwrites and deletions for a defined retention period, and that the protection cannot be removed or shortened by any user, including the account root user. Which S3 feature should the engineer use?
Easy848A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe must replace all null values in a specific column with the string 'UNKNOWN' and then convert the column to uppercase. The engineer wants to apply these steps in a repeatable recipe. Which combination of DataBrew transforms should be used?
Medium849A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing CSV files and writes to an Amazon Redshift table. The job runs successfully but the Redshift table ends up empty. The engineer checks the AWS Glue job run metrics and sees that the job processed 0 rows. The S3 bucket contains files under the prefix 'data/'. The Glue crawler created a table with the correct schema. What is the MOST likely cause of the empty output?
Medium850Refer to the exhibit. A data engineer runs the above AWS CLI command to view the table metadata in the AWS Glue Data Catalog. The data is stored as CSV in S3 with partitions by year and month. When querying the table using Amazon Athena, no data is returned. What is the most likely cause?
Medium851A data engineer is designing a data ingestion pipeline to load data from an on-premises Oracle database into Amazon Redshift. The pipeline must capture changes (inserts, updates, deletes) with low latency and minimal impact on the source database. Which combination of AWS services should the engineer use?
Medium852Refer to the exhibit. A data engineer notices that the Redshift cluster 'mycluster' does not have automated backups beyond 7 days. However, the compliance team requires a minimum of 35 days of backup retention. What should the engineer do?
Medium853A data engineer needs to store streaming data from IoT devices for real-time analytics. The data has a fixed schema and requires low-latency queries. Which AWS service should be used?
Easy854A data engineer is designing a data lake on Amazon S3. The data is frequently accessed by multiple analytics services, and the company needs to enforce fine-grained access control based on data tags. Which combination of AWS services should be used?
Hard855A company uses a Kinesis Data Firehose delivery stream to load data into an S3 bucket. The data is in JSON format and must be converted to Parquet before landing in S3. Which steps are required to achieve this? (Choose THREE.)
Hard856A company runs a Redshift cluster for analytics. The data engineering team notices that COPY commands from S3 are failing for large files (>1 GB) with the error 'S3ServiceException: SlowDown'. What is the most effective solution?
Hard857A data engineer needs to store JSON documents that are frequently accessed by a low-latency web application. The data does not require complex queries, and the access pattern is primarily by a key. Which AWS service is most appropriate?
Easy858A data engineer is managing an AWS Glue Data Catalog that contains metadata for tables in Amazon S3. The security team requires that access to the Data Catalog be restricted based on the user's department, and that users can only see tables that belong to their department. The Data Catalog tables are tagged with a 'Department' key. Which AWS feature should the engineer use to enforce this requirement?
Hard859A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a folder structure. The engineer wants to use AWS Glue crawlers to automatically infer schemas and create tables in the AWS Glue Data Catalog. The crawler should run daily to detect new files and schema changes. Which configuration should the engineer use for the crawler?
Easy860A company uses Amazon DynamoDB with on-demand capacity for a gaming application that experiences unpredictable traffic spikes. The application reads the same set of 'hot' items frequently. Users report high latency during peak hours. Which action would MOST effectively reduce read latency for the hot items?
Hard861A data engineer needs to monitor the performance of an Amazon Redshift cluster. Which Amazon CloudWatch metric should the engineer monitor to detect disk space issues?
Easy862The exhibit shows an S3 bucket policy. What is the effect of this policy?
Medium863A data engineer is configuring a VPC for an Amazon Redshift cluster. The cluster must be accessible only from a specific on-premises network via a Direct Connect connection. Which TWO actions should the engineer take to meet this requirement? (Choose TWO.)
Hard864A data engineer manages an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another bucket. The security team mandates that all data at rest in both buckets be encrypted with customer-managed AWS KMS keys, and that each service use a distinct key. The Glue job's IAM role currently has s3:GetObject and s3:PutObject permissions but jobs fail with an access denied error when writing output. What is the MOST likely cause?
Hard865A data engineer is configuring an AWS Glue crawler to catalog data stored in an Amazon S3 bucket. The security team requires that all data in transit between the crawler and S3 be encrypted using TLS. Which configuration should the engineer implement to meet this requirement?
Easy866A company uses Amazon S3 to store raw data and needs to transform it into Parquet format for analytics. The transformation job runs daily on a schedule. Which AWS service is BEST suited for this task?
Medium867A data engineer runs an AWS Glue job that reads from a JDBC connection to a PostgreSQL database. The job fails with a 'Connection timed out' error. The Glue job runs in a VPC with the appropriate security group. What is the most likely cause?
Medium868Which THREE factors should be considered when choosing between AWS Glue and Amazon EMR for data transformation? (Choose three.)
Hard869A company uses Amazon Kinesis Data Streams to ingest clickstream data from a website. The data is consumed by an AWS Lambda function that writes to Amazon DynamoDB. The Lambda function is seeing high error rates due to DynamoDB write throttling. Which action should be taken to reduce throttling?
Medium870A company is using AWS Glue to run ETL jobs that transform data from Amazon DynamoDB to Amazon S3. The DynamoDB table has a large number of items (over 10 million) and is heavily used by production applications. The Glue job reads the entire DynamoDB table each time it runs, causing increased read capacity consumption and affecting production performance. The team wants to reduce the impact on the source DynamoDB table while still keeping the S3 data up-to-date. What should the team do?
Easy871A data engineer is building a data pipeline that ingests sensitive data into Amazon S3 and then processes it with AWS Glue. The security team requires that the data be encrypted at rest using a customer managed key in AWS KMS, and that the engineer be able to audit all key usage. The engineer creates a KMS customer managed key and configures the S3 bucket to use SSE-KMS with that key. The Glue job's IAM role has been granted kms:Decrypt and kms:GenerateDataKey permissions on the key. However, when the Glue job runs, it fails with an access denied error related to KMS. Which additional action should the engineer take to resolve the error?
Hard872A data engineer is troubleshooting a Kinesis Data Firehose delivery stream that ingests JSON log data from web servers. The stream is configured to transform records with an AWS Lambda function and deliver to an Amazon S3 bucket. Recently, the stream has been failing with 'InvalidData' errors. Which action should the engineer take to resolve the issue?
Easy873A data engineer is using AWS Glue to run an ETL job that reads from Amazon S3, performs a join between two large datasets, and writes the result to Amazon Redshift. The job is taking longer than expected, and the engineer suspects data skew. Which technique can help mitigate data skew in the join?
Hard874A data engineer is using AWS Step Functions to orchestrate a data pipeline that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The engineer needs to ensure that if the AWS Glue job fails, the pipeline retries the job up to three times before failing the entire execution. Which Step Functions state should the engineer use to implement this retry logic?
Medium875A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job must handle upserts (inserts and updates) into an existing Redshift table based on a primary key. The engineer needs to ensure that the job efficiently processes only changed records and minimizes data movement. Which two AWS Glue features or techniques should be used to achieve this? (Choose two.)
Hard876A company is designing a data ingestion pipeline for real-time sensor data from thousands of devices. The data must be processed with low latency and stored in Amazon S3. Which TWO services would be appropriate for this use case? (Choose TWO.)
Medium877A company uses Amazon S3 to store log files. The security team notices that some objects are being accessed from an unexpected AWS account. The data engineer needs to identify which specific IAM user or role is accessing the objects. Which AWS service should be used to get this information?
Medium878A data engineer needs to restrict access to an Amazon S3 bucket so that only objects encrypted with a specific AWS KMS key can be uploaded. Which S3 bucket policy condition should be used?
Easy879Refer to the exhibit. A data engineer runs the AWS CLI command shown to encrypt a file using AWS KMS. The command succeeds. Later, the engineer tries to decrypt the file using the same key but without providing an encryption context. The decryption fails. What is the most likely reason?
Hard880A data engineer needs to troubleshoot why an AWS Glue job is failing with a 'Insufficient Memory' error. The job processes a 10 GB dataset. Which step should the engineer take FIRST?
Easy881A data engineer needs to design a data ingestion pipeline that ingests CSV files from an Amazon S3 bucket, transforms the data by adding a timestamp column, and loads it into an Amazon Redshift table. The pipeline should run automatically whenever a new file is uploaded to the S3 bucket. Which AWS service should be used to trigger the transformation?
Medium882A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer notices that queries are slow and scan large amounts of data. The data is stored in CSV format without compression. Which action should the engineer take to improve query performance and reduce cost?
Easy883A data engineer is using AWS Glue to process a large dataset stored in Amazon S3. The dataset is partitioned by year/month/day and consists of Parquet files. The engineer notices that the Glue job is running slowly and consuming excessive DPU hours. The job performs a join between two large tables and writes the output back to S3. Which optimization technique should the engineer implement to improve performance and reduce cost?
Hard884Refer to the exhibit. A data engineer runs two queries on an Athena table partitioned by 'ds'. Both queries scan the same amount of data. What does this indicate?
Medium885The exhibit shows an IAM policy attached to a role used by an AWS Glue ETL job. The job reads from an S3 bucket and writes to another S3 bucket. However, the job fails with an access denied error when trying to write to the output bucket. What is the most likely cause?
Hard886A data engineer is building a data lake on Amazon S3. The engineer needs to catalog metadata for data stored in Parquet format and make it queryable by Amazon Athena. The data is partitioned by year, month, and day in the S3 path. Which AWS service should the engineer use to create and manage the table definitions and partitions?
Easy887A data engineer is setting up an AWS Glue ETL job that reads data from an Amazon S3 bucket and writes to another S3 bucket. The security team requires that all data in transit be encrypted using TLS. The engineer has configured the job to use the appropriate S3 endpoints. Which additional configuration is necessary to enforce TLS for data in transit between AWS Glue and Amazon S3?
Easy888A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition into Redshift and ensure the load is efficient. Which method should the engineer use?
Medium889Refer to the exhibit. A data engineer applies this S3 bucket policy to an S3 bucket. What is the effect of this policy?
Medium890A company runs an AWS Glue ETL job that reads data from Amazon S3, transforms it, and writes back to S3 in a different partition structure. The job uses the 'spark.sql.shuffle.partitions' option set to 200. After the job completes, the output has many small files. The data engineer wants to minimize the number of output files while maintaining job performance. Which action should the engineer take?
Hard891A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is consumed by an AWS Lambda function that processes each record and writes to an Amazon DynamoDB table. Recently, the Lambda function has been failing with 'ProvisionedThroughputExceededException' from DynamoDB. The Lambda function uses the AWS SDK to batch write items in batches of 25. The DynamoDB table has on-demand capacity mode. The stream has 10 shards, and the Lambda function is configured with a batch size of 100 and 5 concurrent invocations per shard. What step should the team take to resolve the issue?
Hard892A data engineer is troubleshooting an AWS Glue job that reads from Amazon RDS MySQL and writes to Amazon S3. The job runs successfully but takes longer than expected. The engineer wants to optimize performance. Which THREE actions would improve job performance?
Hard893Which TWO AWS services can be used to ingest streaming data from a mobile application into Amazon S3 for near-real-time analytics? (Choose 2.)
Medium894A data engineer is designing a streaming ingestion pipeline using Amazon Kinesis Data Streams. The stream receives records from thousands of IoT devices, and the engineer must ensure that records from the same device are processed in order. The engineer also needs to scale the stream to handle peak loads without manual intervention. Which two actions should the engineer take? (Choose two.)
Hard895A retail company uses Amazon Redshift for its data warehouse. The security team requires that all data in the cluster be encrypted at rest using a hardware security module (HSM) to manage the encryption keys. The data engineer needs to configure the Redshift cluster accordingly. Which action should the data engineer take?
Easy896A company stores time-series sensor data in Amazon S3. They need to query the data using SQL with minimal latency and no infrastructure management. Which service should they use?
Easy897A data engineer is troubleshooting an AWS Glue job that is reading from an Amazon Kinesis Data Stream. The job is configured with a 1-minute window and is supposed to process the latest records. However, the engineer notices that the job is reprocessing old data from the stream. What is the most likely cause of this issue?
Medium898A data engineer is designing an Amazon S3 data lake and needs to enforce schema-on-read for a dataset that is queried by Amazon Athena. The data is stored as Parquet files partitioned by year, month, and day. The engineer wants to minimize the amount of data scanned by queries that filter on a specific date range. Which approach should the engineer take?
Medium899A data engineer needs to store semi-structured JSON transaction logs for analytics. The logs are written once and rarely accessed. The storage must be cost-effective. Which AWS service should be used?
Easy900A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources in Parquet format, and the schema evolves over time. Which approach allows querying the data with Amazon Athena while supporting schema evolution?
Medium901A company needs to centralize audit logs from multiple AWS accounts into a single S3 bucket. Which service should be used to aggregate these logs?
Easy902A company runs an Apache Spark job on Amazon EMR that writes output to an S3 bucket. The job fails with the error 'S3AccessDeniedException' when writing the final output, but earlier stages succeed. The EMR cluster uses a service role and an instance profile. The S3 bucket policy allows access from the VPC only. What is the MOST likely cause?
Hard903A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs and Amazon EMR steps. The state machine has a task that starts a Glue job and waits for completion using the 'StartJobRun' API with the 'sync' integration. Occasionally, the Step Functions execution fails with the error: 'States.TaskFailed: Glue job failed with error: ResourceNumberLimitExceededException'. The engineer confirms that the Glue job itself runs successfully when triggered manually. What is the most likely cause of this intermittent failure?
Hard904A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon S3 in Parquet format. The source data is in JSON format and contains nested structures. The engineer needs to flatten the nested data and write it to Parquet. Which AWS Glue transform should the engineer use to flatten the nested structure?
Medium905A data engineer runs an AWS Glue ETL job that joins a 4 TB Parquet dataset in Amazon S3 with a small 40 MB reference lookup table stored as a single CSV file in S3. The join is taking hours and the job frequently fails with executor out-of-memory errors. The engineer wants to reduce shuffle and memory pressure with the least development effort. Which approach should the engineer take?
Medium906A data engineer needs to schedule a recurring AWS Glue ETL job to run every day at 2:00 AM UTC. The job must be triggered automatically without manual intervention. Which AWS service should the engineer use to create this schedule?
Easy907A data engineer needs to ensure that data in transit between an Amazon RDS for PostgreSQL database and an application is encrypted. Which configuration should be used?
Easy908A company wants to enable automatic encryption for all new objects written to an S3 bucket. The bucket has existing objects that are unencrypted. Which solution meets these requirements with the least operational overhead?
Medium909A data engineer needs to automate the backup of an Amazon RDS for PostgreSQL database. Which AWS service can be used to schedule and manage the backups?
Easy910A company uses AWS DMS to replicate data from an on-premises Oracle database to Amazon RDS for MySQL. The full load completes successfully, but ongoing replication (CDC) is failing with a 'Failed to add supplemental logging' error. What should the data engineer do to resolve this issue?
Easy911Refer to the exhibit. An S3 event notification is configured to trigger an AWS Lambda function when objects are created in 'my-bucket'. The Lambda function processes the JSON file and writes results to Amazon DynamoDB. The function fails with a timeout error. Which action should the engineer take to resolve the issue?
Easy912A data engineer is designing a pipeline to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real-time with minimal latency, and the engineer wants to use a fully managed service that automatically scales. The data is in JSON format and needs to be converted to Parquet before storage. Which AWS service should be used to achieve this?
Medium913A data engineer is using AWS Lake Formation to manage access to a data lake stored in Amazon S3. The engineer grants a data analyst SELECT permission on a table in the AWS Glue Data Catalog. However, when the analyst queries the table using Amazon Athena, they receive an error that they are not authorized to access the underlying S3 data. What is the MOST likely reason?
Hard914A company has an Amazon RDS for MySQL DB instance with read replicas. The primary DB instance fails. What is the correct procedure to promote a read replica to become the new primary?
Medium915A data engineering team is using Amazon EMR to process large datasets stored in Amazon S3. The cluster uses Spot Instances for cost savings. During processing, the team notices that tasks are failing due to Spot Instance interruptions. The team needs to make the EMR job resilient to Spot interruptions without increasing costs significantly. Which solution should they implement?
Medium916A data engineer is using AWS Lake Formation to manage access to a data lake stored in Amazon S3. The engineer needs to grant a data analyst read access to specific columns in a table registered in the AWS Glue Data Catalog, while hiding other columns that contain personally identifiable information. The analyst uses Amazon Athena to query the table. Which Lake Formation feature should the engineer use?
Hard917Refer to the exhibit. The S3 bucket policy above is applied to the bucket "example-bucket". An IAM user attempts to upload an object to the bucket without specifying any encryption header. What is the outcome?
Medium918A company is migrating an on-premises Apache Cassandra database to Amazon Keyspaces. The database has a table with a partition key of 'user_id' and a clustering column of 'timestamp'. The application frequently queries the last 10 records for a given user. Which table design in Keyspaces would provide the BEST query performance for this access pattern?
Medium919A data engineer sees this AWS Glue table definition in the Data Catalog. The engineer wants to query this table with Amazon Athena, but the query returns zero rows. What is the MOST likely cause?
Medium920A company is using Amazon Athena to query data stored in S3. Queries are failing with 'HIVE_INVALID_PARTITION' errors. What is the most likely cause?
Medium921A data engineer needs to ingest JSON files from an S3 bucket into a DynamoDB table. The files are updated hourly and contain new records. Which AWS service should be used to trigger a Lambda function for each new object?
Easy922A company uses Amazon S3 to store sensitive data. The security team wants to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted at rest using server-side encryption with AWS KMS managed keys (SSE-KMS). Which bucket policy statement should be added to enforce this requirement?
Medium923A data engineer needs to ensure that data stored in Amazon S3 is automatically deleted after 30 days. Which S3 feature should be used?
Easy924A data engineer is troubleshooting an AWS Glue ETL job that fails with the error: 'An error occurred while calling o137.pyWriteDynamicFrame. No such file or directory: s3://bucket/output/part-00000.parquet'. The job reads from a JDBC source and writes to S3. What is the most likely cause?
Easy925A company is using Amazon DynamoDB for a gaming application. They want to store player session data that expires after 24 hours. Which DynamoDB feature should be used?
Easy926Which TWO actions can reduce the cost of an Amazon S3 bucket that stores infrequently accessed data? (Choose 2.)
Medium927A data engineer is using AWS Glue Studio to create an ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The source data is in JSON format and contains nested structures. The engineer needs to flatten the nested structures and write the output in Parquet format. The job must be efficient and scalable. Which transformation should the engineer use in Glue Studio to flatten the nested data?
Hard928A data engineer is setting up Amazon S3 event notifications to trigger an AWS Lambda function when new objects are uploaded. Which TWO actions are required to enable this?
Easy929A data engineer is troubleshooting a failed AWS Glue ETL job that reads from an S3 bucket and writes to an Amazon Redshift table. The job fails with a permission error. Which IAM policy addition is MOST likely required for the Glue job's role?
Medium930A data engineer is using AWS Database Migration Service (AWS DMS) to migrate a 4 TB on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must complete in a single maintenance window, and the source database cannot be taken offline for more than 30 minutes. The engineer configures a full load plus change data capture (CDC) task. During testing, the full load phase takes 14 hours. Which configuration change will most effectively reduce the time required for the full load phase?
Medium931A data engineer needs to ingest data from a SaaS application (Salesforce) into Amazon S3 on a daily basis. Which TWO AWS services can be used for this purpose? (Choose TWO.)
Easy932A data engineer is managing an Amazon Redshift cluster and needs to load data from Amazon S3. The engineer uses the COPY command but encounters the error: 'S3ServiceException: Access Denied.' The Redshift cluster has an IAM role attached with permissions to access the S3 bucket. What is the most likely cause of this error?
Medium933A data engineer must choose a storage service for a new application that requires single-digit millisecond latency at any scale, a flexible schema, and automatic scaling of throughput without provisioning capacity. The access pattern is key-value lookups by user ID with occasional range queries on a sort key. Which AWS service should the engineer select?
Easy934A healthcare company stores patient records in an S3 bucket encrypted with SSE-S3. The data engineering team uses AWS Glue ETL jobs to process this data and load it into an Amazon Redshift cluster for analytics. Recently, the security team mandated that all sensitive data must be encrypted at rest using customer-managed keys (CMK) in AWS KMS, and that the keys must be rotated automatically every year. The team updated the S3 bucket to use SSE-KMS with a CMK and enabled automatic key rotation. However, after the change, the Glue ETL jobs that read from the S3 bucket started failing with 'Access Denied' errors. The Glue job uses an IAM role named 'GlueETLRole' that has the following permissions: s3:GetObject on the bucket, kms:Decrypt and kms:GenerateDataKey on the CMK, and all necessary Glue permissions. The Redshift cluster is also encrypted with a different CMK, and the Glue role has kms:Decrypt on that key as well. What is the most likely cause of the failure?
Hard935A data engineer manages an Amazon Redshift cluster that experiences performance degradation during complex analytical queries. The engineer notices that some queries spill to disk. The engineer wants to improve query performance by optimizing the distribution style and sort keys. Which action should the engineer take first?
Hard936A company needs to automate the detection of sensitive data in Amazon S3 and generate reports. Which AWS service should be used?
Medium937A data engineer maintains an AWS Glue Data Catalog with databases for several lines of business. Auditors require that every change to table definitions, partition additions, and schema edits in the catalog be recorded with the identity of the caller and that the records be retained for 365 days in a dedicated S3 bucket. Which solution should the data engineer implement?
Hard938Refer to the exhibit. A data engineer has configured an S3 event notification to send an event to an SQS queue when objects are created in the 'incoming/' prefix. The engineer wants to trigger an AWS Lambda function to process the object. However, the Lambda function is not being invoked. What is the most likely cause?
Hard939A data engineer needs to back up an Amazon DynamoDB table daily. The backup must be restorable to a specific point in time within the last 24 hours. Which solution meets these requirements with the LEAST operational overhead?
Easy940Refer to the exhibit. A data engineer runs this AWS CLI command to execute an Athena query. What is the purpose of the EncryptionConfiguration parameter?
Medium941A data engineer maintains an AWS Glue ETL job that reads JSON from Amazon S3 and writes Parquet to a second bucket. Downstream consumers report that numeric fields occasionally arrive as strings and timestamps are sometimes null. The engineer must make the job resilient to these schema variations without failing the run. Which approach should the engineer take?
Hard942A data engineer manages an Amazon Redshift cluster that runs a nightly ETL load followed by complex analytical queries. Users report that queries during the day are slower than expected, and the team wants to isolate the ETL workload so it cannot consume resources needed by the analytical queries. The cluster uses provisioned nodes. What is the MOST appropriate solution?
Medium943A media company stores video files in an S3 bucket. The files are processed by a fleet of EC2 instances that read the files, add watermarks, and write the output back to the same bucket. Recently, the processing jobs have been failing with '500 Internal Server Error' and '503 Slow Down' errors. The data engineer checks the S3 bucket metrics and sees that the PUT/GET request rate is consistently above 5,500 requests per second for a single prefix. The engineer needs to resolve the errors with minimal changes to the application code. Which course of action should the engineer take?
Medium944A data engineer is troubleshooting an AWS Glue job that reads from an Apache Kafka topic using a Glue connector. The job fails with 'TimeoutException'. The Kafka cluster is in a VPC. Which step should the engineer take FIRST?
Easy945A data engineer notices that an AWS Glue ETL job that processes streaming data from Amazon Kinesis Data Streams is failing intermittently with a 'ResourceNotFoundException' error for the Kinesis stream. The job has been running successfully for weeks. Which action should the engineer take to resolve the issue?
Medium946A data engineer is using AWS Glue Studio to build a job that reads from an Amazon Kinesis Data Stream, performs a 5-minute tumbling window aggregation, and writes results to Amazon S3. The job must run continuously and handle late-arriving records within the window. Which configuration should the engineer use?
Hard947A data engineer needs to run a Python-based transformation on each object as it lands in an Amazon S3 bucket. The objects are small (under 10 MB), arrive sporadically, and must be processed within seconds. Which approach is MOST appropriate?
Easy948A company is using AWS Glue to process data stored in Amazon S3. The Glue job runs successfully but takes longer than expected. Which TWO actions can reduce the job runtime?
Easy949A data engineer runs a weekly AWS Glue ETL job that processes data from Amazon DynamoDB to Amazon S3. The job reads the entire table every time, which is slow and expensive. The job needs to process only items that changed since the last run. Which solution should the engineer implement?
Hard950A data engineer is building a data lake on Amazon S3 and must enforce that all objects containing personally identifiable information are encrypted with a customer managed AWS KMS key, while allowing automatic key rotation and audit of key usage. Objects must remain readable by an AWS Glue job and an Amazon Athena workgroup. Which configuration should the engineer choose?
Medium951A data engineer needs to encrypt data in transit between an Amazon RDS for MySQL instance and an application. Which solution should be used?
Easy952A data engineer needs to transform JSON data into Parquet format using AWS Glue. The input data has nested fields. Which Glue feature should be used to flatten the nested structure?
Medium953A data engineering team uses Amazon S3 to store raw data files. They have an AWS Glue ETL job that reads from an S3 bucket, transforms the data, and writes to a Redshift cluster. The job runs daily and has been failing intermittently with the error: 'An error occurred while calling o143.pyWriteDynamicFrame. S3 Access Denied'. The team has confirmed that the IAM role used by the Glue job has s3:GetObject and s3:PutObject permissions on the bucket and all objects. The Redshift cluster is in the same VPC and the Glue connection is configured correctly. What is the most likely cause of the failure?
Medium954A data engineer is using AWS Database Migration Service (AWS DMS) to migrate a large on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must minimize downtime, so the engineer needs to capture ongoing changes while the initial full load runs. Which AWS DMS task configuration should the engineer use?
Medium955A data engineer manages an AWS Glue ETL job that writes Parquet files to Amazon S3. Downstream Amazon Athena queries started returning duplicate rows after the job was modified to enable job bookmarks. The job reads from an S3 source prefix where new files are appended hourly and the transformation includes a join that reorders records. Which action will most reliably eliminate the duplicate rows while preserving incremental processing?
Medium956A data engineer needs to schedule an AWS Glue ETL job to run every hour and process new data that arrives in an S3 bucket. The job should only process files that have been added since the last run. Which approach should the engineer use to track which files have been processed?
Easy957A company wants to schedule a nightly batch job to copy data from an on-premises PostgreSQL database to Amazon S3. The solution must minimize operational overhead. Which AWS service should be used?
Easy958A company uses Amazon Redshift for data warehousing. The data engineering team notices that queries are slow due to high disk I/O. The team wants to improve query performance without changing the cluster configuration. Which action should the team take?
Medium959A financial services company needs to share sensitive customer data with a third-party analytics firm. The data resides in an S3 bucket encrypted with an AWS KMS customer managed key. The third party has their own AWS account. Which combination of steps is required to securely share the data? (Choose TWO.)
Hard960A company has an AWS Glue ETL job that reads from an RDS MySQL instance and writes to S3. The security team requires that the connection to RDS be encrypted and that credentials be rotated automatically. Which configuration should be used?
Hard961A data engineer is using Amazon Athena to query Parquet data in Amazon S3. Queries are slow and scan more data than expected. The data is partitioned by year/month/day in S3, but the AWS Glue Data Catalog table has no partition metadata. Which action will improve query performance and reduce data scanned?
Hard962A data engineer manages an AWS Glue ETL job that reads JSON files from Amazon S3 and writes to Amazon Redshift. The job recently started failing with the error: 'Unable to find catalog table' when trying to access a table in the AWS Glue Data Catalog. The engineer confirms that the table exists in the Data Catalog and that the IAM role used by the job has glue:GetTable permissions. What is the most likely cause of this error?
Medium963A data engineer is troubleshooting an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The job runs successfully but 5% of records are missing after the load. The engineer suspects data consistency issues. Which THREE actions could help diagnose and resolve the problem? (Choose THREE.)
Hard964A company uses Amazon S3 to store large datasets. The data engineering team needs to provide access to specific objects in the bucket to external partners using presigned URLs. Each URL should expire after 12 hours. The team wants to ensure that the presigned URLs cannot be used to access other objects in the bucket. Which approach should be taken?
Hard965A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to an Amazon Redshift cluster. The security team requires that the data be encrypted in transit between Glue and Redshift. Which configuration should the engineer implement to meet this requirement?
Medium966A data engineer needs to store JSON documents that are accessed by a serverless application using AWS Lambda. The documents are frequently updated and need low latency (single-digit milliseconds) for read and write operations. Which AWS service should the engineer use?
Easy967A data engineer is using AWS Glue ETL to transform a large dataset in S3. The job processes 2 TB of data daily and currently runs for 6 hours. The engineer wants to reduce runtime without changing the transformation logic. What is the best approach?
Medium968A data engineer needs to ingest data from an on-premises Oracle database into Amazon S3. The data volume is about 500 GB initially, with daily incremental updates of 10 GB. The pipeline must minimize operational overhead. Which AWS service should be used for the initial and incremental loads?
Easy969A data engineer is monitoring an Amazon Kinesis Data Analytics application that processes real-time clickstream data. The application uses a Flink application with multiple operators. The engineer notices that the 'millisBehindLatest' metric is increasing steadily. Which action is MOST likely to reduce the lag?
Hard970A data engineer is using Amazon EMR to process large datasets. The cluster uses a mix of Spot Instances and On-Demand Instances. The engineer wants to reduce costs while ensuring the job can complete even if Spot Instances are reclaimed. Which TWO actions should the engineer take? (Choose two.)
Medium971A data pipeline uses AWS DMS to replicate data from an on-premises Oracle database to Amazon S3 in Parquet format. The pipeline has been running successfully for months, but recently the DMS task status shows 'failed' with the error: 'The source database is running out of archive log space.' Which action should the engineer take to prevent this error?
Hard972A data engineer is troubleshooting a nightly ETL job that reads data from an RDS MySQL instance and writes to an S3 bucket in Parquet format. The job runs on an EMR cluster and uses PySpark. Recently, the job started failing with 'OutOfMemoryError' in the executor logs. The data volume has grown 30% in the last month. Which is the MOST efficient solution to resolve this issue without changing the code?
Medium973Which THREE factors should be considered when choosing between Amazon Kinesis Data Streams and Amazon Kinesis Data Firehose for real-time data ingestion? (Choose three.)
Hard974A company wants to ingest real-time streaming data from thousands of IoT devices into AWS for immediate processing. Which service is designed for ingesting large volumes of streaming data with low latency?
Easy975A company runs a critical PostgreSQL database on Amazon RDS. The database experiences high read latency during peak hours. The data engineer needs to reduce read latency with minimal changes to the application. Which solution is MOST effective?
Hard976A data engineer is troubleshooting an ETL job that reads from an S3 bucket encrypted with SSE-KMS. The job is failing with an error indicating that the IAM role does not have permission to decrypt the data. What is the most likely missing permission?
Hard977A company wants to audit all changes to IAM policies in their AWS account. Which AWS service should be used to record these changes for compliance purposes?
Easy978A data engineer is using AWS Lake Formation to manage fine-grained access to a data lake in Amazon S3. The engineer grants a data analyst SELECT permission on a table but wants to ensure that the analyst cannot access columns containing sensitive data such as social security numbers. The table is registered in the AWS Glue Data Catalog. Which Lake Formation feature should the engineer use to restrict access to specific columns?
Medium979A data engineer needs to audit all changes to IAM policies in an AWS account. Which AWS service should be used?
Easy980A company uses AWS Glue to run ETL jobs that process data from Amazon S3 and load into Amazon Redshift. The jobs have recently started failing with 'Out of Memory' errors. The data volume has increased 3x in the past month. Which is the MOST effective solution to resolve this issue without redesigning the job?
Hard981A data engineer needs to load data from an on-premises Oracle database to Amazon S3 daily. The table is 500 GB and grows by 50 MB per day. The load must capture only new and changed rows since the last run. Which solution is MOST cost-effective and requires the least maintenance?
Medium982Refer to the exhibit. A CloudFormation stack outputs the Glue job name and S3 bucket names. The Glue job transforms CSV files from the raw bucket to Parquet in the processed bucket. However, the Glue job is failing with an error that it cannot write to the processed bucket. What is the most likely cause?
Hard983A company uses Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data must be transformed in real-time using a custom Lambda function. Which TWO steps are required to enable this? (Choose TWO)
Easy984A healthcare company stores patient records in an Amazon S3 bucket and uses AWS Lake Formation to manage access for multiple analytics teams. The compliance team requires that any column containing patient identifiers be masked by default for all users except a privileged data steward role. Which Lake Formation feature should the data engineer implement to meet this requirement?
Medium985A company uses Amazon Redshift for analytics. The data engineering team wants to improve query performance for frequently used aggregate queries. Which TWO actions would help achieve this?
Medium986A data engineer is using Amazon Redshift and needs to improve query performance for a large fact table that is frequently joined with a much smaller dimension table. The engineer wants to minimize data movement during joins. Which distribution style should be used for the dimension table?
Medium987A data engineer is building an AWS Glue ETL job that reads records from an Amazon Kinesis Data Stream and writes them to Amazon S3 in Parquet format. The job must checkpoint its progress so that it can resume without reprocessing data after a failure. Which AWS Glue mechanism should the engineer configure to track the stream position?
Medium988A company uses Amazon S3 to store large CSV files and runs Amazon Athena queries on them. The queries are becoming slower as data grows. A data engineer suggests converting the files to Apache Parquet format and partitioning the data. What is the primary benefit of converting to Parquet?
Medium989A company is building a data lake on AWS and must encrypt data at rest. Which services can provide server-side encryption for data stored in Amazon S3? (Choose TWO.)
Medium990A company ingests JSON data from an S3 bucket into a Glue ETL job. The data contains nested structures and arrays. The team wants to flatten the data into a tabular format for analysis in Athena. Which Glue transformation is appropriate?
Hard991A data engineer must give an Amazon Redshift cluster the ability to load data from an Amazon S3 bucket using the COPY command. The security team prohibits embedding long-term AWS credentials in SQL and requires that access be revoked automatically when the cluster is deleted. The S3 bucket is encrypted with SSE-KMS using a customer managed key. Which approach should the data engineer use?
Medium992A data engineer notices that an Amazon Redshift cluster’s storage usage is increasing rapidly due to many UPDATE and DELETE operations. The engineer needs to reclaim storage space and improve query performance. Which action should be taken?
Hard993A data engineer needs to store semi-structured data (JSON logs) from thousands of IoT devices. The data must be schema-less, highly scalable, and support low-latency queries by device ID and timestamp. Which AWS service should the engineer use?
Easy994A company stores application logs in an Amazon S3 bucket. A compliance policy states that log objects must be retained for exactly 90 days and then permanently deleted, and that no one, including administrators, should be able to delete them earlier. The data engineer must enforce this with the least effort. What should the engineer do?
Easy995A company runs a Redshift cluster and notices that query performance has degraded over time. The data engineer suspects that table statistics are stale. What should the engineer do to improve query performance?
Medium996A company runs a daily batch ETL job using AWS Glue. The job processes 500 GB of data from Amazon RDS to Amazon S3. The job currently uses a single DPU and takes 6 hours to complete. The team wants to reduce runtime to under 1 hour without increasing costs significantly. Which approach should they use?
Hard997A data engineer is configuring an Amazon S3 bucket to store sensitive financial data. The company requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption key be automatically rotated every year. The engineer creates a KMS customer managed key and enables automatic rotation. When uploading objects using the AWS CLI, the engineer uses the --sse aws:kms parameter but does not specify a key ID. What is the result of this configuration?
Hard998A company uses AWS DMS to replicate data from an Amazon RDS for MySQL database to Amazon S3. Which TWO configurations are required to enable continuous change data capture (CDC) from MySQL?
Hard999Which THREE factors should be considered when choosing a partition key for an Amazon DynamoDB table?
Hard1000A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to an S3 bucket. The data is delivered in 5-minute intervals. However, the engineer notices that the data in S3 is often delayed by up to 30 minutes. Which configuration change would most likely reduce the delay?
Hard1001A data engineer manages an Amazon S3 data lake with millions of small JSON files ingested continuously. Amazon Athena queries over this data are slow and expensive because each query scans many small objects. The engineer wants to improve query performance and reduce cost without changing the data content. Which solution should the engineer implement?
Medium1002A data engineer must mask the last four digits of a credit card column in an Amazon Redshift table so that analysts in a specific role see masked values while a fraud team sees the full values. The engineer wants a solution that applies to all queries without modifying each analyst's SQL. Which approach should the engineer use?
Medium1003A data engineer is configuring an AWS Glue crawler against an Amazon S3 path that contains CSV files with inconsistent column counts across files. The crawler keeps creating multiple tables for the same data and the engineer wants a single table with a merged schema. Which crawler configuration should the engineer change?
Medium1004A company needs to transform JSON data from an S3 bucket into a structured format for Amazon Redshift. The transformation should be done serverlessly. Which service should be used?
Easy1005A data engineer maintains an AWS Glue job that incrementally processes new files in Amazon S3 using job bookmarks. After a schema change in the source data added a new column, the engineer updated the Glue Data Catalog table. Subsequent job runs still process only previously seen files and ignore newly arrived objects. The engineer verifies that new files exist in the prefix and that the bookmark state was not reset. Which factor most likely explains why new files are being skipped?
Hard1006A financial services company uses AWS Glue ETL jobs to process sensitive customer data stored in Amazon S3. The data is encrypted at rest with SSE-KMS using a customer-managed key. Recently, the security team discovered that the Glue job's IAM role has an overly permissive policy that allows the 'kms:Decrypt' action for all KMS keys in the account. The company wants to follow the principle of least privilege. The Glue job runs on a schedule and reads from a specific S3 bucket. The security team needs to update the IAM policy to restrict KMS decryption to only the specific key used for that bucket. What should they do?
Medium1007A data engineer needs to allow an IAM user to rotate the secret in AWS Secrets Manager for an RDS database. Which IAM action should be included in the policy?
Medium1008A data engineer needs to store semi-structured JSON data from IoT devices. The data is written frequently and read occasionally. Which AWS service is MOST cost-effective for this use case?
Easy1009A data engineer needs to ensure that an Amazon S3 bucket containing sensitive customer data is encrypted at rest. Which AWS service can be used to manage the encryption keys?
Easy1010A data engineer notices that an AWS Glue ETL job is failing with an OutOfMemory error when processing a large dataset. The job uses a Standard worker type. Which action is MOST effective to resolve this issue without changing the job script?
Medium1011A data engineer is using AWS Glue to process data stored in Amazon S3. The engineer needs to ensure that the AWS Glue job can access the S3 bucket securely without hardcoding credentials. Which approach should the engineer use?
Medium1012A data engineer is designing a multi-region disaster recovery solution for Amazon RDS for PostgreSQL. The primary region must have a standby in a different Availability Zone, and the secondary region must have a readable replica that can be promoted in case of failure. Which configuration meets these requirements?
Hard1013A company uses Amazon S3 to store raw data and AWS Glue to run ETL jobs. The data is partitioned by date in the format 'year=YYYY/month=MM/day=DD'. A new data source started sending data with a different date format 'YYYY-MM-DD'. The Glue crawler is configured to create a single table for the entire bucket. The crawler runs daily, but it is not detecting the new partitions from the new data source. The existing partitions are in the format 'year=2024/month=05/day=10', while the new data is stored as '2024-05-10/' without the key-value structure. How should the engineer modify the data pipeline to include the new data?
Easy1014A company is using AWS Lake Formation to manage permissions on a data lake. Which of the following are valid ways to grant access to a user or role? (Choose THREE.)
Medium1015A data engineer is troubleshooting an Amazon Redshift cluster that is not responding to queries. The engineer suspects that the cluster may have been accidentally deleted. Which AWS service should be used to investigate the deletion?
Medium1016A company uses Amazon DynamoDB for a gaming application. The application experiences throttling during peak hours. The table's read and write capacity is provisioned. Which TWO actions can reduce throttling?
Medium1017A company is using Amazon Kinesis Data Streams to ingest real-time clickstream data from a website. The data is consumed by an Amazon Kinesis Data Analytics for Apache Flink application that performs real-time analytics. The Flink application writes its results to an Amazon S3 bucket. The company has noticed that the Flink application is experiencing high checkpoint failure rates, causing delays. The CloudWatch metrics show that the checkpoint size is large and increasing. The data engineer needs to reduce the checkpoint size. Which action should the data engineer take?
Medium1018A data engineer runs an AWS Glue Studio job that reads JSON from Amazon S3 and writes to a partitioned Parquet table. The job currently runs for six hours. Profiling shows that a small number of partitions contain millions of rows while most contain a few hundred. Which change will MOST improve runtime?
Hard1019A media company stores millions of thumbnail images in an Amazon S3 bucket. Analysts run ad hoc queries against the image metadata, which is kept as JSON objects in the same bucket. Query latency is unpredictable and costs are rising because Athena scans large volumes of JSON for every query. The team wants faster queries and lower scan cost while keeping the data in S3 and queryable with SQL. Which change should the data engineer make?
Medium1020A data engineer is monitoring an Amazon Kinesis Data Stream and notices that the 'WriteProvisionedThroughputExceeded' metric is frequently elevated. The stream has 5 shards and is used by multiple producers. What is the BEST action to resolve this issue?
Medium1021A data engineer needs to ensure that an AWS Glue job has access to an Amazon RDS database in a private subnet. The Glue job will run in a VPC and requires a security group and subnet configuration. Which combination of steps should the engineer take?
Easy1022A data engineering team is troubleshooting a failing AWS Glue ETL job that processes data from an S3 bucket. The job writes output to another S3 bucket. The job fails with an AccessDenied error when writing to the output bucket. The IAM role used by the job has the following policy attached: {"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":["s3:GetObject","s3:ListBucket"],"Resource":["arn:aws:s3:::input-bucket/*","arn:aws:s3:::input-bucket"]}]}. What is the most likely cause of the failure?
Medium1023A data engineer is building an AWS Glue Studio visual ETL job that reads JSON files from Amazon S3, applies a transformation, and writes to Amazon Redshift. During a test run, the job fails with an error indicating that the dynamic frame could not be written because the target table schema does not match the incoming data. The engineer needs to ensure the job automatically reconciles schema differences such as missing columns and data type mismatches during the write. Which action should the engineer take?
Medium1024A data engineer is building an AWS Glue ETL job that reads from an AWS Glue Data Catalog table backed by Amazon S3. The job must process only records added since the last successful run to reduce cost and runtime. The source data is partitioned by year, month, and day. Which approach should the engineer use?
Hard1025A data engineer is designing an ingestion pipeline that uses AWS Glue to read from an Amazon RDS for PostgreSQL database. The job must read only rows changed since the previous run and must not scan the entire table each night. The source table has a last_updated timestamp column that is updated on every write. (Choose two.)
Medium1026A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The security team requires that data be encrypted at rest using a customer-managed AWS KMS key, and that the Glue job be able to decrypt the source data and encrypt the target data. The engineer has already created a KMS key and attached a key policy that allows the Glue service role to use the key for encrypt and decrypt operations. However, when the job runs, it fails with an access denied error related to KMS. What is the most likely cause of the failure?
Medium1027A company is using AWS Glue to run ETL jobs that transform data from S3 to Redshift. The jobs are failing intermittently with out-of-memory errors. Which THREE actions can help resolve this issue? (Choose THREE.)
Medium1028A company has a 100 TB dataset stored on-premises in a Hadoop cluster. They want to ingest this data into Amazon S3 for processing with AWS Glue. The company has a limited time window and a slow internet connection. Which strategy is MOST appropriate?
Hard1029A company runs an Amazon EMR cluster with Spark jobs that process data from Amazon S3. The data engineer receives an alert that one of the Spark jobs failed with an OutOfMemoryError. The job processes large files and uses the default Spark configurations. Which configuration change is MOST likely to resolve the issue?
Medium1030A company uses AWS Glue to process data from multiple sources. The data is stored in an Amazon S3 data lake. The company needs to transform the data using a custom Python library that is not available in the default Glue environment. What is the MOST efficient way to make this library available to the Glue jobs?
Medium1031A data engineering team needs to transform CSV files stored in Amazon S3 into Parquet format using AWS Glue. The files are partitioned by date and are updated hourly. Which AWS Glue feature should be used to automatically detect the schema and partition structure?
Easy1032A data engineer is deploying an Amazon Redshift cluster that must be accessible only from within a private VPC and must not have a public IP address. The cluster will be queried by an Amazon EMR cluster in the same VPC and by on-premises BI tools over a VPN connection. Which configuration should the engineer choose?
Medium1033A data engineer is configuring an Amazon S3 bucket that stores sensitive customer records for analytics. The security team requires that all data be encrypted at rest with keys that are rotated automatically every year and that access be auditable per key. The engineer must minimize operational overhead. Which encryption configuration should be used?
Medium1034A company uses Amazon DynamoDB to store user session data. The table has a partition key of user_id and a sort key of session_start. The workload is read-heavy and eventually consistent reads are acceptable. The table is provisioned with 1000 RCUs and 500 WCUs. During peak hours, the application experiences throttling on read operations, but CloudWatch shows that the consumed read capacity is well below the provisioned amount. What is the most likely cause of the throttling?
Hard1035A data pipeline ingests JSON data from an S3 bucket using AWS Glue. The JSON files contain nested structures, and the team wants to flatten them for analysis in Amazon Athena. Which Glue transformation is most appropriate?
Hard1036Which TWO are valid approaches to troubleshoot a slow Amazon Redshift query? (Choose two.)
Hard1037A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. The state machine includes a task that runs an AWS Glue job and then waits for its completion. The engineer notices that the Step Functions execution times out after 15 minutes, even though the Glue job takes about 30 minutes to complete. The Step Functions state machine has a timeout of 1 hour. What is the most likely cause of the timeout?
Hard1038A data engineer manages an AWS Glue job that reads from an Amazon S3 bucket containing PII. The security team requires that the data be encrypted at rest using a customer-managed AWS KMS key, and that the engineer be able to audit key usage. The engineer has already created a KMS key. Which combination of steps should the engineer take to meet these requirements?
Medium1039A data engineer needs to store JSON documents that are frequently read and written by a web application. The data has a flexible schema and requires low-latency queries on primary key lookups. Which AWS service is MOST suitable?
Easy1040A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. One of the Glue jobs occasionally fails due to transient network issues. The engineer wants the Step Function to retry the failed Glue job up to 3 times with exponential backoff before failing the entire workflow. Which Step Functions state configuration should be used?
Hard1041A company needs to ingest streaming data from thousands of IoT devices into Amazon S3 for long-term storage and analytics. The data arrives continuously at a rate of 5 MB per second and must be stored in a compressed format to reduce storage costs. The solution should be highly available and require minimal management. Which AWS service should the company use?
Easy1042A data engineer needs to ingest streaming data from thousands of devices sending JSON messages via HTTP POST. The data should be stored in Amazon S3 with minimal latency and also be available for real-time analytics. Which combination of services is MOST appropriate?
Medium1043A data engineer needs to store and analyze time-series data from IoT devices. The data volume is 10 GB per day, and the queries are mostly on the most recent 7 days of data. The engineer wants to minimize storage costs while retaining historical data for 1 year. Which combination of AWS services is most cost-effective?
Medium1044A company is using Amazon S3 for data lake storage. They need to query the data directly using SQL without loading it into a database. Which AWS service should be used?
Easy1045A company uses AWS Glue to transform data stored in S3. The Glue job runs daily and processes data in the range of hundreds of GB. The data engineer wants to optimize the job for cost and performance. Which THREE actions should be taken? (Choose THREE.)
Hard1046A company wants to use Amazon Redshift Spectrum to query data in Amazon S3. The data is in Parquet format and partitioned by date. Which step is required to enable Redshift Spectrum?
Easy1047A data engineer needs to share an S3 bucket with another AWS account. They want to ensure that the objects in the bucket remain encrypted with SSE-KMS using a customer managed key. What additional step is required for cross-account access?
Medium1048Which TWO options are valid methods to ingest on-premises relational database data into Amazon S3 for analytics? (Choose 2.)
Medium1049A data engineer is using AWS Lake Formation to manage access to a data lake in Amazon S3. The company wants to grant a data analyst read-only access to specific columns in a table stored in the AWS Glue Data Catalog. The analyst should not be able to see other columns or any rows that contain sensitive data. The engineer sets up Lake Formation permissions on the table, granting SELECT on specific columns. However, when the analyst queries the table using Amazon Athena, they can see all columns. What is the most likely reason?
Medium1050A company needs to store files that are accessed by multiple EC2 instances in a VPC. The files must be concurrently accessible and durable. Which storage solution should the data engineer choose?
Easy1051The exhibit shows an AWS CLI command and its output. A data engineer wants to copy only objects larger than 10 MB from the S3 bucket to another bucket for processing. Which approach should be used to automate this task?
Medium1052A data engineer manages an Amazon DynamoDB table for order events. Reads and writes are evenly spread across a partition key with very high cardinality, but during flash sales the table throttles with ProvisionedThroughputExceededException even though consumed capacity is below the provisioned total. Which cause is MOST likely?
Hard1053A data engineer needs to ingest streaming data from thousands of IoT devices into AWS for real-time processing. The data volume peaks at 5 GB/min. Which AWS service should be used as the ingestion endpoint?
Easy1054A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON and writes to a partitioned Parquet table. The engineer wants to reduce job cost and improve read performance. (Choose two.)
Medium1055A company is building a data lake on Amazon S3 and needs to ingest data from multiple sources. Which of the following AWS services can be used to ingest and transform data in near real-time? (Select TWO.)
Medium1056A data engineer needs to run a transformation on streaming data using SQL-like queries without managing servers, and the output must be written to Amazon S3 in near real time. The source is an Amazon Kinesis Data Stream. Which AWS service is the MOST appropriate to perform the transformation?
Easy1057A data engineering team is using AWS Glue to catalog data in an S3 data lake. They have a Glue crawler that runs daily to update the Data Catalog. Recently, they noticed that the crawler is taking longer to run and sometimes fails because of a timeout. The team suspects the issue is due to the large number of small files in the S3 bucket. They need to improve crawler performance and reliability. Which solution should they implement?
Easy1058A data engineer is setting up Amazon S3 bucket policies for a data lake. Which TWO statements are true regarding S3 bucket policies? (Choose TWO.)
Easy1059A company wants to ingest streaming data from Apache Kafka into Amazon S3 for long-term storage and analytics. The data is in JSON format and must be delivered to S3 with minimal effort and no custom code. Which AWS service should the data engineer use?
Easy1060A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is then processed by a Kinesis Data Analytics application running SQL queries. The analytics application is falling behind and processing records with increasing latency. The stream has 4 shards, and the average record size is 5 KB. What is the MOST effective way to improve processing latency?
Hard1061A data engineer is troubleshooting a failed AWS Glue job that writes results to Amazon S3. The error log shows 'AccessDenied' when trying to list the bucket. Which IAM policy statement should the engineer add to the Glue job's role?
Easy1062A data engineer is optimizing an AWS Glue ETL job that reads from Amazon S3 and writes to Amazon Redshift. The job currently uses a single large file and takes hours to complete. The engineer wants to improve performance by using partitioning and parallelism. Which TWO actions should the engineer take? (Choose two.)
Hard1063A data engineer needs to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket for long-term storage. The data is in JSON format and must be delivered within 60 seconds of arrival. The engineer wants a fully managed solution that requires minimal code. Which service should the engineer use?
Easy1064A company uses Amazon S3 as its data lake. A data engineer needs to enforce encryption of data at rest using server-side encryption with AWS KMS. Which S3 bucket property should be configured?
Easy1065Refer to the exhibit. A data engineer creates an Amazon Redshift table with the above DDL. The engineer runs a query to find all orders for a specific customer within a date range. Which statement about query performance is correct?
Easy1066A data engineer is optimizing an Amazon RDS for MySQL database that experiences high write throughput. The engineer wants to improve write performance and reduce latency. Which TWO database-level configuration changes can help achieve this?
Medium1067A data engineer is using AWS Glue to process a large dataset where a small number of partitions contain disproportionately more rows than others, causing some executors to run much longer than others and the job to take hours. The engineer wants to redistribute the data across partitions before a join operation to improve performance. Which technique should the engineer apply?
Hard1068A company is ingesting real-time clickstream data into Amazon S3 using Amazon Kinesis Data Firehose. The data is semi-structured and the company wants to transform the data into Parquet format and partition it by year, month, day, and hour. Which TWO steps should be taken to achieve this? (Choose TWO.)
Medium1069A data engineer runs an AWS Glue job that writes Parquet files to Amazon S3. The job frequently fails with an error indicating too many small files are being written, causing slow downstream Athena queries. The engineer wants to reduce the number of output files without changing the transformation logic. Which action should the engineer take?
Hard1070Which THREE factors should a data engineer consider when choosing between AWS Glue and Amazon EMR for a data transformation job? (Choose three.)
Hard1071A data engineer needs to run a transformation in AWS Glue where each record must be processed independently and the output schema is known ahead of time. The transformation should operate on a DynamicFrame and return a DynamicFrame. Which Glue transform is designed for this row-by-row operation?
Easy1072A data engineer is using AWS Glue Studio to create a job that joins data from two Amazon S3 sources: a large fact table and a small dimension table. The job performs a join and then writes the result to Amazon S3 in Parquet format. The engineer notices that the job is running slowly and consuming many DPUs. Which optimization technique should the engineer apply to improve performance?
Hard1073A company has an Amazon DynamoDB table with a provisioned write capacity of 1000 WCU. During a flash sale, the write traffic spikes to 5000 WCU for 10 minutes. The table is not auto-scaled. Which action should the data engineer take to handle the spike without throttling?
Hard1074A company uses AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration is taking longer than expected. The task status shows 'Full load in progress' with a low 'Table throughput (rows/s)'. Which action would MOST improve throughput?
Hard1075A company uses AWS Glue ETL to transform data from Amazon RDS for MySQL to Amazon S3. The Glue job reads from a JDBC connection. The job runs once daily and processes all records, but the data volume is growing. Which change would improve performance and reduce costs?
Medium1076A data engineer needs to run a transformation on a large dataset stored in Amazon S3 using AWS Glue Studio. The transformation is a simple column rename and filter that can be expressed visually. The engineer wants to minimize development time and avoid writing PySpark code. Which approach should the engineer use?
Easy1077A company needs to ingest real-time clickstream data from a web application into Amazon S3 for analytics. The data must be available within minutes of generation. Which AWS service should be used to capture and deliver this streaming data?
Easy1078A data pipeline using AWS Glue jobs is failing with 'Insufficient capacity' errors for Spark executors. Which action should the data engineer take to resolve this?
Medium1079A company is experiencing high costs from Amazon Redshift. The data engineer wants to optimize costs. Which THREE actions should the engineer take? (Choose THREE.)
Hard1080A data engineer runs an AWS Glue ETL job that reads semi-structured JSON from Amazon S3, flattens nested arrays, and writes Parquet to a partitioned S3 location. Job runs are becoming expensive because Glue reprocesses all historical partitions on every run. The engineer wants subsequent runs to process only newly arrived data. Which approach should the engineer take with the LEAST operational overhead?
Medium1081A data engineer runs an Apache Spark application on Amazon EMR that writes partitioned Parquet output to Amazon S3. A downstream AWS Glue crawler registers the table in the Data Catalog, but Athena queries return zero rows for partitions added by the most recent run, while older partitions query correctly. The S3 objects exist and are readable. Which action should the engineer take?
Hard1082A data engineer is designing a streaming ingestion pipeline using Amazon Kinesis Data Streams. The stream has 10 shards, and the data volume is expected to grow by 50% over the next month. The engineer needs to ensure that the pipeline can scale without manual intervention. Which approach should be used?
Hard1083A company wants to migrate on-premises data to Amazon S3 using AWS DataSync. The data is stored on an NFS file server and the total volume is 50 TB. The network bandwidth between the on-premises data center and AWS is 1 Gbps (gigabit per second). What is the primary factor that will determine the total time required for the initial data transfer?
Easy1084A company uses Amazon S3 to store raw data and runs AWS Glue ETL jobs to transform it into Parquet. The data is then queried using Amazon Athena. Queries are slow and expensive due to high scan volumes. Which THREE design changes can improve query performance and reduce costs? (Select THREE.)
Medium1085A data engineer runs the command shown to check the encryption configuration of an S3 bucket. The output shows SSEAlgorithm: AES256. What does this mean?
Easy1086A company is using Amazon Athena to query data in an S3 bucket. Queries are failing with the error 'HIVE_PATH_ALREADY_EXISTS'. The data is partitioned by year, month, day. What is the MOST likely cause?
Medium1087Refer to the exhibit. A data engineer queries AWS CloudTrail to investigate a PutObject event. What does the exhibit reveal about the object sensitive.csv?
Medium1088A data engineer is migrating a large Oracle data warehouse to Amazon Redshift. The engineer needs to ensure optimal performance. Which TWO practices should the engineer follow?
Medium1089A data engineer is building an AWS Glue ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3. The DynamoDB table has a large number of items, and the engineer needs to ensure the job reads the data efficiently without consuming too much provisioned throughput. Which method should the engineer use to read from DynamoDB?
Medium1090A data engineer is configuring an S3 bucket for storing sensitive customer data. The bucket must be encrypted at rest using an AWS Key Management Service (KMS) key that is managed by the data engineering team. The team wants to ensure that only users with explicit permission can decrypt the data. Which S3 encryption option should be used?
Medium1091A data engineer needs to ingest log files from multiple EC2 instances into Amazon S3. The logs are written to local disk on each instance. The engineer wants a simple agent-based solution that can collect, compress, and upload logs to S3 with minimal configuration. The solution must support incremental uploads (only new log lines) and handle log rotation. What should the engineer use?
Easy1092A retail company uses Amazon DynamoDB to store product catalog data. The table has a partition key of ProductID and a sort key of Category. The company needs to retrieve all products in a specific category, sorted by ProductID. Which operation should be used?
Easy1093A data engineer is using AWS Glue DataBrew to profile a dataset stored in Amazon S3. The profile shows that a column named country contains values such as 'US', 'usa', 'United States', and 'U.S.A.' The engineer needs to standardize these values to a single canonical form before loading the data into Amazon Redshift. Which DataBrew transformation should the engineer apply?
Medium1094A data engineer is designing a data lake on S3 with sensitive data. The security policy mandates that data must be encrypted at rest and in transit, and that an inventory of all objects must be maintained for compliance. Which actions should be taken?
Medium1095Which THREE storage classes in Amazon S3 are designed for infrequently accessed data with millisecond retrieval times? (Select THREE.)
Medium1096A company is using Amazon RDS for MySQL with Multi-AZ deployment. The primary DB instance experiences a hardware failure, causing automatic failover to the standby. After the failover, the application reports that the database endpoint is unreachable for about 60 seconds. What is the MOST likely cause?
Medium1097A data engineer maintains an AWS Glue ETL job that processes JSON files from Amazon S3 and writes Parquet to another S3 location. The job has been running successfully for months. Recently, the job started failing intermittently with the error 'Unable to infer schema for JSON'. The engineer confirms the source bucket contains valid JSON files. Which action should the engineer take to resolve the failure?
Medium1098A data engineer needs to ingest streaming data from an IoT fleet into Amazon S3 for near-real-time analytics. The data volume is approximately 5 GB per hour, and each event is less than 1 KB. Which AWS service should be used as the ingestion endpoint?
Easy1099A data engineer needs to ingest data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real time, and the engineer wants to minimize operational overhead by using a fully managed service that can also transform the data format from JSON to Parquet. Which AWS service should be used?
Easy1100A company has an S3 bucket that stores logs for compliance. The compliance team requires that objects are retained for 7 years and cannot be deleted or overwritten. Which S3 feature should be used?
Easy1101A data engineer needs to store large amounts of data that is accessed infrequently but must be retrieved immediately when needed. Which Amazon S3 storage class is most cost-effective?
Easy1102A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3. The data is stored as CSV files. The downstream team requires the data to be in Apache Parquet format. Which change should the data engineer make to the DMS task?
Easy1103A data engineer is building an AWS Glue job that reads from a JDBC source and must retrieve the database password at runtime without hardcoding it in the script or job parameters in plaintext. The company already stores the password in AWS Secrets Manager. Which action should the engineer take?
Easy1104A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 source with many small JSON files and writes to Amazon S3 in Parquet. The job runs slowly and produces many tiny output files. The engineer wants to improve throughput and reduce the number of output files without changing the source data layout. (Choose two.)
Hard1105A company is using Amazon DynamoDB as a data store for a real-time application. The application reads a single item by primary key and occasionally updates it. The data engineer notices high read latency during peak hours. Which TWO actions would most effectively reduce read latency?
Medium1106A company uses Amazon RDS for PostgreSQL to store customer data. The data engineer needs to ensure that the database can be restored to any point in time within the last 35 days. The engineer also wants to minimize the impact on the production database during backups. What should the engineer do?
Easy1107A data engineer maintains an Amazon Kinesis Data Streams pipeline that feeds an AWS Lambda consumer. During traffic spikes, the Lambda function is throttled and records are reprocessed, causing duplicate entries in the downstream Amazon S3 sink. The engineer needs to reduce duplicates with the LEAST code change. What should the engineer do?
Medium1108A company is designing a data lake on AWS and must comply with GDPR requirements. The company needs to implement data masking for personally identifiable information (PII) columns in Amazon Redshift. Which feature should be used?
Medium1109A data engineer is troubleshooting a step function that orchestrates ETL jobs. The state machine fails with 'State Machine Execution Throttled' error. What should the engineer do to resolve this?
Medium1110A company uses AWS Kinesis Data Streams to ingest real-time data. The data engineer notices that the stream's 'WriteProvisionedThroughputExceeded' error occurs frequently during peaks. Which action should be taken to resolve this issue?
Medium1111A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift table on a daily schedule. The data is in CSV format and the schema matches. Which service is simplest for this batch ingestion?
Easy1112A data engineer is troubleshooting an issue where an IAM role used by AWS Glue cannot read data from an S3 bucket encrypted with SSE-KMS. The bucket policy allows the role to perform s3:GetObject. What additional permission is needed?
Hard1113A data engineer needs to transfer 50 TB of historical data from an on-premises HDFS cluster to Amazon S3. The on-premises network has a 1 Gbps link to AWS. The transfer must complete within 5 days. Which solution is MOST cost-effective and meets the requirements?
Medium1114A company uses Amazon DynamoDB with global tables in three AWS Regions. The data engineer needs to ensure that writes to the table in us-east-1 are replicated to other regions with minimal latency. Which DynamoDB feature should be used?
Medium1115A data engineer needs to ingest streaming data from an Amazon Kinesis Data Stream into an Amazon S3 bucket. The data must be delivered in near real-time and stored in Parquet format for efficient querying. The engineer wants to minimize custom code. Which solution should the engineer use?
Easy1116A data engineer maintains an AWS Glue ETL job that reads JSON files from Amazon S3, applies transformations using a DynamicFrame, and writes Parquet to another S3 bucket. The job currently reads the entire source prefix on every run, but the source is now partitioned by year/month/day and only new partitions need processing. The engineer wants to process only new partitions without changing the job's transformation logic. Which change should the engineer make?
Hard1117A data engineer must load data from an Amazon DynamoDB table into an Amazon S3 data lake nightly. The table is approximately 800 GB and the nightly window is tight. The engineer wants a fully managed, serverless option that exports the table to S3 without consuming DynamoDB read capacity or writing custom code. Which solution meets these requirements?
Easy1118A company has a DynamoDB table with a partition key of 'user_id' and a sort key of 'timestamp'. They need to query all items for a user within a date range. Which query operation should be used?
Hard1119A data engineer is building an AWS Glue ETL job that reads a large Amazon S3 dataset of nested JSON files and must flatten the nested arrays into separate rows for downstream analytics. The engineer needs the most efficient, code-free way to apply this transformation within the Glue job. Which approach should the engineer use?
Medium1120A data engineering team is troubleshooting a slow AWS Glue ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3 in Parquet format. The job processes 50 GB of data. Which action would most effectively improve job performance?
Hard1121A company uses AWS Lake Formation to manage fine-grained access to a data lake in Amazon S3. A data analyst needs to query a table in the AWS Glue Data Catalog that contains columns with sensitive data. The analyst must be able to see only non-sensitive columns and only rows where the region column equals 'US'. The analyst uses Amazon Athena for queries. Which Lake Formation permission model should the data engineer implement?
Hard1122A company wants to ingest streaming data from IoT devices into Amazon S3 using Amazon Kinesis Data Firehose. The data must be transformed from JSON to Parquet format before landing in S3. What is the SIMPLEST way to achieve this?
Easy1123A data engineer is configuring an S3 bucket policy to allow cross-account access for a partner account to read objects. The bucket is encrypted with SSE-KMS using a customer-managed key. What additional configuration is needed to allow the partner account to decrypt the objects?
Medium1124A data engineer is using AWS Glue Studio to build an ETL job that reads semi-structured JSON from Amazon S3. The source files contain nested arrays and inconsistent keys, and the engineer wants the job to automatically infer the schema at runtime without a Data Catalog table. Which transform or configuration should the engineer use to read the data most reliably?
Medium1125A company uses Amazon S3 to store sensitive data. The data engineer needs to ensure that all data in transit between the S3 bucket and clients is encrypted. Which configuration should the engineer implement?
Medium1126A company runs an Amazon Redshift cluster with 10 RA3 nodes. The data warehouse stores 50 TB of data. The company notices that queries are slow and the cluster's storage utilization is high. The data engineer needs to improve query performance and reduce storage costs without changing the cluster's node count. Which action should the engineer take?
Hard1127A data engineer needs to run a one-time AWS Glue ETL job that reads data from an Amazon S3 bucket and writes transformed Parquet files to another S3 bucket. The job does not need a schedule and should be run immediately after creation. Which method should the engineer use to start the job?
Easy1128A data engineer is designing a data lake on Amazon S3 with sensitive data. The engineer needs to ensure that data at rest is encrypted and that access is logged for compliance. Which TWO actions should the engineer take? (Choose TWO.)
Hard1129A company needs to ingest data from multiple SaaS applications (e.g., Salesforce, Marketo) into Amazon S3 for analytics. The data sources have different schemas and update frequencies. Which AWS service should be used to build this ingestion pipeline with minimal code?
Medium1130A company uses an Amazon RDS for MySQL DB instance with Multi-AZ deployment. The primary DB instance fails unexpectedly. What happens to the database endpoint?
Easy1131A company wants to ingest data from multiple SaaS applications into Amazon S3 using a fully managed service that supports schema discovery and transformation. Which AWS service should they use?
Easy1132A data engineer needs to move 50 TB of existing data from an on-premises data center into Amazon S3 as a one-time migration. The data center has a 1 Gbps internet connection that is shared with production traffic, and the migration must complete within two weeks without disrupting production. Which approach should the engineer use?
Easy1133Refer to the exhibit. A data engineer configured the lifecycle policy shown. The 'logs/' prefix contains important audit logs. After 365 days, what happens to the objects?
Medium1134Which TWO statements about Amazon Redshift data distribution are correct? (Choose two.)
Easy1135A company's Amazon Redshift cluster is running slowly. The data engineer suspects that table design is the cause. Which TWO design practices can improve query performance? (Choose TWO.)
Medium1136A company stores application logs in Amazon S3 in JSON format. The logs are partitioned by year/month/day. A data engineer needs to create a table in the AWS Glue Data Catalog so that Amazon Athena can query the logs efficiently. The engineer wants to minimize query costs and ensure that new partitions are automatically recognized. Which combination of actions should the engineer take?
Medium1137A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job uses the write_dynamic_frame.from_jdbc_conf method. The engineer notices that the job is slow and sometimes fails due to connection timeouts. Which action should the engineer take to improve performance and reliability?
Hard1138A data engineer needs to ensure that all objects written to an S3 bucket are encrypted with SSE-KMS using a specific customer managed key, and that any upload without that encryption is rejected. The engineer has created the bucket and the KMS key. Which approach will enforce this requirement at the bucket level?
Medium1139A company uses AWS Glue ETL jobs to transform data stored in Amazon S3. The job reads data in Parquet format, applies transformations, and writes the output back to S3 in Parquet format. The team wants to improve the job's performance and reduce costs. Which action is MOST effective?
Easy1140A company must comply with a regulation that requires logging all access to sensitive data stored in Amazon S3. Which AWS services can be used to capture and store access logs? (Choose TWO.)
Easy1141A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes transformed records to an Amazon Redshift cluster. The job must load data in parallel and use an Amazon Redshift IAM role for authentication. Which connection option should the engineer configure in the Glue job to enable parallel loading and IAM-based authentication?
Medium1142A data engineer manages an Amazon DynamoDB table used for a high-traffic gaming leaderboard. The table uses on-demand capacity mode and has a partition key of UserId (string) with no sort key. The leaderboard must retrieve the top 100 scores across all users. Currently, the engineer scans the entire table and sorts the results in application code, which takes several seconds and consumes large amounts of read capacity. What should the engineer do to improve the performance of retrieving the top scores?
Medium1143A company uses Amazon DynamoDB for a gaming leaderboard. The table has a partition key of 'GameId' and a sort key of 'Score'. The application needs to query the top 10 scores for a given game. Which DynamoDB feature should be used for optimal performance?
Hard1144Refer to the exhibit. A data engineer runs an AWS Glue ETL job that writes output to an S3 bucket. The job fails with the error shown. What is the most likely cause?
Hard1145A company uses AWS Glue to run ETL jobs that process data from an Amazon RDS for MySQL database and load it into an Amazon S3 data lake. The Glue job runs daily and processes incremental data. Recently, the job has been taking longer than expected. The engineer checks the CloudWatch logs and sees that the job is spending most of its time on the 'Reading from JDBC' phase. The MySQL table has 10 million rows and is indexed on the primary key. The Glue job uses a 'job bookmark' to track processed data. The engineer wants to improve the performance of the read phase. Which action is most likely to help?
Easy1146A company uses Amazon QuickSight for data visualization. The data engineer needs to ensure that users can only see data relevant to their department. The data is stored in Amazon S3 and is accessed via SPICE. The engineer has created datasets in QuickSight and wants to implement row-level security (RLS). The dataset contains a column 'Department' that indicates which department a row belongs to. The engineer has configured RLS rules using a separate permissions dataset. However, users report that they can see all rows, not just their department's rows. What is the most likely reason?
Easy1147Refer to the exhibit. An AWS Glue job is failing with 'AccessDenied' when trying to write to the 'data-lake-bucket' which is encrypted with an AWS KMS key. The IAM role used by the Glue job has the attached policy shown. What is the MOST likely cause of the failure?
Hard1148A company uses Amazon RDS for MySQL with Multi-AZ deployment. The primary instance fails, and automatic failover occurs. After failover, the application experiences higher latency. What is the most likely cause?
Medium1149A data engineer is using AWS Glue Studio to build a job that reads from Amazon S3, applies a filter, and writes to Amazon Redshift. The engineer needs to ensure the job can be rerun safely without creating duplicate rows in Redshift if a previous run partially succeeded. Which design choice best meets this requirement?
Hard1150A company wants to ingest streaming data from thousands of IoT devices into AWS for real-time analytics. Which AWS service is best suited for this purpose?
Easy1151Order the steps to set up a Kinesis Data Analytics application for real-time stream processing.
Medium1152A data engineer needs to ingest data from a relational database (MySQL) into Amazon S3 for analytics. The database is 500 GB and the job must run daily with incremental updates. Which AWS service is BEST suited for this task?
Easy1153A company needs to store relational data that requires complex joins and transactional consistency. The workload is predictable and the data size is less than 500 GB. Which AWS service is MOST cost-effective for this use case?
Easy1154A data engineer must give an AWS Glue ETL job temporary access to data in an Amazon S3 bucket without creating long-term IAM user access keys. The job runs on a schedule and must retrieve credentials automatically. Which mechanism should the engineer use?
Easy1155A company runs a daily batch processing job on Amazon EMR that reads data from Amazon S3 and writes results back to S3. The job takes longer than expected. The engineer wants to monitor the job's resource utilization. Which AWS service should be used to collect and visualize metrics such as CPU and memory usage of the EMR cluster's nodes?
Easy1156A data engineer is using AWS Step Functions to orchestrate a data pipeline that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The engineer needs to ensure that if the Glue job fails, the pipeline stops and does not proceed to the EMR step. Which Step Functions state type should be used to handle the error and stop the execution?
Easy1157A company uses Amazon Redshift for its data warehouse. During a routine audit, the data engineer discovers that some queries are returning stale data even though the underlying source data has been updated. The engineer confirms that the COPY command completes successfully and that no errors are reported. Which action should the engineer take to ensure queries reflect the latest data?
Hard1158A data engineer runs an AWS Glue job that reads a large partitioned Parquet dataset from Amazon S3 and writes aggregated results to another S3 prefix. The job runs daily and currently reprocesses the entire dataset each time, which is becoming expensive. The engineer wants subsequent runs to process only data added since the last successful run, based on the job's state. Which feature should the engineer enable?
Hard1159A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The data engineer needs to ensure minimal downtime during migration. Which AWS service should be used to facilitate the migration?
Easy1160A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data must be transformed in real-time using custom Python code before being stored in Amazon S3. Which AWS service should be used to perform this transformation?
Medium1161A company needs to ingest data from multiple SaaS applications (Salesforce, Marketo) into Amazon S3 for analytics. The data volume is moderate (~100 GB per day). The pipeline must handle schema changes, deduplicate records, and provide low latency (under 1 hour). Which THREE services should be used? (Choose THREE.)
Hard1162A company is migrating a legacy on-premises ETL pipeline to AWS. The pipeline processes daily batch files from an FTP server. The data must be transformed using complex business logic before being loaded into Amazon Redshift. Which THREE AWS services should be used for this migration?
Hard1163A data engineer is designing a data lake on Amazon S3. The data lake will store raw data, transformed data, and curated datasets. The engineer needs to ensure that raw data is immutable (never overwritten or deleted) and that only authorized users can access the transformed data. Which combination of S3 features should the engineer use?
Easy1164A data engineer needs to ingest data from multiple SaaS applications (Salesforce, Marketo) into Amazon S3 for a data lake. The data volumes are moderate and the sync needs to be scheduled daily. Which AWS service is most appropriate for this task?
Easy1165Match each AWS data migration tool to its primary function.
Medium1166A data engineer is designing a data warehouse on Amazon Redshift. The workload includes many ad-hoc queries that filter on a high-cardinality column, such as customer_id, and join large dimension tables. The engineer wants to improve query performance by choosing an appropriate distribution style and sort key. Which combination should the engineer use?
Hard1167A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files and must flatten the nested structures before writing to Amazon Redshift. The job uses the Glue DynamicFrame API. The engineer wants the transformation to run as a single pass without an intermediate shuffle. Which operation should the engineer use?
Hard1168A data engineer is configuring an Amazon Redshift cluster for a workload that runs large nightly ELT jobs loading data from Amazon S3 and then executes complex analytical queries. The team wants to improve query performance and reduce the time the cluster spends on data loading. Which TWO configuration choices should the engineer make? (Choose two.)
Medium1169A data engineer runs an AWS Glue for Apache Spark job that writes partitioned Parquet output to Amazon S3. The engineer notices thousands of small output files in each partition, which slows downstream Amazon Athena queries. The job reads from a large S3 source and uses default partitioning. Which change should the engineer make to reduce the number of small files?
Medium1170A company is streaming IoT sensor data to Amazon Kinesis Data Streams. The data is JSON with a schema that changes occasionally. They want to load the data into Amazon S3 in Parquet format partitioned by date and sensor_id. Which approach is MOST cost-effective and operationally efficient?
Medium1171A data engineer is using Amazon Athena to query data stored in Amazon S3. The data is partitioned by date, but the engineer notices that queries are scanning the entire bucket instead of only the relevant partitions. The table is defined in the AWS Glue Data Catalog. Which action should the engineer take to ensure that Athena only scans the necessary partitions?
Medium1172Refer to the exhibit. A data engineer is troubleshooting an AWS Lambda function that reads from an S3 bucket and writes to a Kinesis Data Stream. The Lambda function fails with an AccessDeniedException when calling the kinesis:PutRecords API. Which change is needed to the IAM policy?
Medium1173A data engineer is using Amazon Kinesis Data Streams to ingest real-time data. The engineer needs to ensure that records are delivered to consumers in the same order they were written and that each record is processed exactly once. Which combination of Kinesis Data Streams features should the engineer use?
Easy1174A company wants to ensure that all S3 buckets are encrypted using server-side encryption. Which AWS service can be used to automatically remediate non-compliant buckets?
Easy1175A data engineer is setting up an Amazon DynamoDB table to store user session data. The table must handle sudden spikes in read traffic during peak hours, and the engineer wants to minimize operational overhead while ensuring consistent performance. The table's read capacity mode should automatically adjust to traffic changes. Which capacity mode should the engineer choose?
Easy1176A data engineer is configuring AWS Glue jobs to access data stored in Amazon S3. The data is encrypted using server-side encryption with AWS KMS (SSE-KMS). The Glue job needs to read and write data to the S3 bucket. Which IAM policy statement should be added to the Glue job's IAM role to allow it to use the KMS key?
Easy1177A company wants to use AWS Glue to transform data stored in Amazon S3. The data is partitioned by date and includes both CSV and Parquet files. The transformation should be optimized for cost and performance. Which THREE actions should the data engineer take? (Choose THREE.)
Medium1178Which TWO statements are true about Amazon Redshift distribution styles? (Choose TWO.)
Medium1179A media company is building a data pipeline to ingest user activity logs from multiple sources into Amazon S3. The logs are JSON files generated every minute. The company wants to use Amazon Athena to query the logs with minimal latency and cost. The current approach is to use AWS Kinesis Data Firehose to deliver the logs to S3 with a prefix like 'logs/2024/01/01/00/file.json'. However, when running Athena queries, the team notices high query costs because Athena scans all files in the 'logs/' prefix even when querying for a specific date. What should the team do to reduce the amount of data scanned by Athena?
Easy1180A data engineer maintains an AWS Glue Data Catalog table for an S3-based dataset. After new files were added, queries in Amazon Athena fail with the error 'HIVE_BAD_DATA: Error parsing field value for field 3: For input string: "N/A"'. The column is defined as bigint in the Data Catalog but contains the string 'N/A' in some records. Which action should the data engineer take to allow Athena to query the data without changing the underlying files?
Medium1181A company runs a nightly ETL job using AWS Glue. The job reads data from a JDBC connection to an on-premises MySQL database. The job fails with an error indicating that the connection pool is exhausted. What is the most likely cause and solution?
Medium1182A company uses AWS Glue ETL jobs to transform data in S3. The job runs successfully but takes longer than expected. The data is in Parquet format and partitioned by date. Which change would most improve performance without increasing cost?
Medium1183A media company stores video files in an Amazon S3 bucket. The bucket policy allows access only from a specific VPC. The company has enabled S3 Server Access Logs to monitor access. Recently, the security team found that some requests were coming from an IP address outside the allowed VPC. They suspect that the bucket policy may have an incorrect condition. What should they check first?
Easy1184A data engineer is configuring an AWS Glue ETL job to read from an Amazon S3 bucket that contains Apache Parquet files partitioned by year, month, and day. The engineer wants the job to only process data for the year 2023 and month 10, and to minimize the amount of data scanned. The Glue job uses the Glue Data Catalog table `sales_data` with the correct partition structure. What is the MOST efficient way to configure the job to read only the required partitions?
Medium1185A data engineer needs to catalog a growing S3 data lake. New CSV files land in s3://analytics/raw/orders/ with a partition structure year=YYYY/month=MM/day=DD/. The engineer must create an AWS Glue Data Catalog table that automatically recognizes these partitions and requires no crawler runs for future dates. Which approach meets these requirements?
Medium1186A data engineer needs to store JSON documents that are frequently updated and require ACID transactions. Which AWS database service is most appropriate?
Easy1187A data engineer is designing an ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver streaming records into an Amazon S3 bucket. The records arrive as JSON, and downstream consumers require Parquet with a stable schema. The engineer must configure the Firehose delivery stream so records are converted to Parquet before landing in S3. (Choose two.)
Medium1188A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB and has a 24/7 uptime requirement. The migration must minimize downtime. Which approach should be used?
Medium1189A data engineer applies the following IAM policy to an IAM user: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::example-bucket/*", "Condition": { "StringEquals": { "s3:x-amz-server-side-encryption": "AES256" } } } ] } ``` The user attempts to download an object from the bucket 'example-bucket' that is encrypted with SSE-S3 (AES256). Will the request succeed?
Medium1190A company uses AWS Glue to process data from Amazon S3. The data contains personally identifiable information (PII). The data engineer needs to automatically detect and mask PII fields before the data is loaded into Amazon Redshift. Which combination of AWS services should be used?
Hard1191A company stores sensitive data in Amazon S3 and needs to ensure that data is encrypted at rest. The security team requires that the company manage its own encryption keys and have the ability to audit key usage. Which S3 encryption option should the data engineer choose?
Easy1192A company needs to store JSON documents that are frequently read and written by a web application. The data must be highly available and durable across multiple Availability Zones. Which AWS database service meets these requirements?
Easy1193A company runs a MySQL database on Amazon RDS. The database size is 500 GB and is experiencing high read traffic. The team wants to improve read performance with minimal operational overhead. Which action should they take?
Easy1194A data engineer needs to transform JSON data from Amazon S3 into Parquet using AWS Glue. The JSON is nested and contains arrays. The engineer wants to flatten the nested structure and write the result to S3 partitioned by a 'region' field. Which combination of Glue transforms should the engineer use?
Medium1195A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Aurora MySQL. The migration is successful, but the ongoing replication task is experiencing high latency. Which configuration change is most likely to reduce latency?
Medium1196A data engineer is troubleshooting a Lambda function that reads from the Kinesis stream 'my-data-stream'. The Lambda function is able to read data but occasionally fails with 'KMS.AccessDeniedException'. What is the most likely cause?
Medium1197A data engineer is monitoring an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The job runs daily and recently started taking significantly longer to complete. The engineer checks the job metrics and notices that the number of DPUs used is consistently at the maximum allocated, and the job's Spark UI shows many tasks spilling to disk. Which action should the engineer take to improve performance?
Medium1198A data engineer must load a 2 GB uncompressed CSV file from Amazon S3 into Amazon Redshift using the COPY command. The cluster is a two-node ra3.xlplus cluster, and the load is running far slower than expected. The engineer wants the fastest reliable improvement without changing the cluster. What should the engineer do?
Easy1199A company uses Amazon Athena to query data in S3. Recently, queries have become slow. The data is stored as CSV files in a partitioned table. What is the most effective way to improve query performance?
Easy1200A data engineer is using AWS Glue DataBrew to profile a dataset in Amazon S3. The dataset contains a column 'customer_id' that should be unique. The engineer runs a profile job and notices that the 'customer_id' column has a uniqueness metric of 98%. The engineer needs to identify the duplicate values. Which DataBrew feature should the engineer use to display the duplicate values?
Hard1201A data engineer notices that an Amazon Redshift cluster is running low on disk space. The cluster has three nodes of type dc2.large. Which action will increase the available storage capacity?
Medium1202A company's security policy states that no S3 bucket in the data platform account may ever be made public, even accidentally. A data engineer must implement a guardrail that blocks any attempt to set a public bucket ACL or public bucket policy, regardless of who makes the change. Which solution enforces this requirement?
Easy1203A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer notices that some JSON records have missing fields and inconsistent schemas. Which AWS Glue feature should be used to handle these inconsistencies and ensure the job does not fail?
Hard1204A company is ingesting streaming data into Kinesis Data Streams. The consumer application experiences high latency due to a single shard bottleneck. What is the most effective way to reduce latency?
Medium1205A data engineer is designing a data lake on Amazon S3. Data is ingested from multiple sources in JSON format. The engineer needs to optimize query performance for Amazon Athena while minimizing storage costs. Which storage strategy should the engineer use?
Medium1206A data engineer maintains an AWS Glue ETL job that reads from an Amazon Kinesis Data Streams stream and writes to Amazon S3 in Parquet format. The job has been running successfully, but after the data volume increased threefold, the job now fails with an error stating that the Glue job's bookmarks are not advancing and the job is reprocessing old data. The engineer has enabled job bookmarks with the default settings. Which action should the engineer take to resolve the issue?
Medium1207A company is ingesting streaming data from social media feeds using Amazon Kinesis Data Streams. The data volume peaks at 10,000 records per second, and each record is up to 1 KB. The company needs to archive the raw data in Amazon S3 in near real-time and also make it available for real-time analytics using Amazon Kinesis Data Analytics. What is the MOST efficient architecture to meet these requirements?
Hard1208Which TWO AWS services can be used to automatically back up an Amazon RDS for SQL Server DB instance? (Choose TWO.)
Easy1209A company uses Amazon Kinesis Data Streams to ingest real-time financial data. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key, and that the key be rotated annually. The data engineer needs to configure the Kinesis stream to meet these requirements. Which combination of actions should the data engineer take?
Medium1210A data engineer has an AWS Glue job that reads JSON files from Amazon S3, applies transformations, and writes Parquet files to another S3 location. The job runs daily and takes about 2 hours. Recently, the job has been failing intermittently with an error indicating that the job bookmark is not being updated correctly, causing duplicate processing of some files. The engineer needs to ensure that only new files are processed on each run. Which action should the engineer take to resolve this issue?
Medium1211A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by year, month, and day. The engineer needs to ensure that AWS Glue crawlers correctly identify the partitions and that Amazon Athena queries can efficiently prune partitions. Which action should the engineer take?
Medium1212A data engineer needs to move data from an Amazon S3 bucket to an Amazon Redshift cluster on a daily schedule. The data is in CSV format and the target table already exists. Which AWS service should the engineer use to automate this task?
Easy1213A data pipeline ingests streaming data from thousands of IoT devices into Kinesis Data Streams. The data must be transformed using a simple field mapping before being stored in S3. Which service should be used to perform the transformation with minimal operational overhead?
Easy1214A company is using Amazon S3 to store critical data and needs to ensure that objects are automatically transitioned to S3 Glacier Deep Archive after 180 days to reduce costs. Which S3 lifecycle action should be configured?
Easy1215A company needs to store archival logs that must be retained for 10 years. The logs are accessed infrequently, but when accessed, retrieval must occur within 12 hours. Which storage class is MOST cost-effective?
Easy1216A company wants to ingest data from SaaS applications (e.g., Salesforce, Marketo) into Amazon S3 for analytics. The data volume is moderate and updates occur frequently. Which AWS service is BEST suited for this task?
Medium1217A data pipeline ingests daily CSV files from an FTP server into an Amazon S3 bucket. The files must be converted to Parquet format and partitioned by date for efficient querying using Amazon Athena. Which AWS service is most suitable for this transformation?
Easy1218A data engineer is designing a pipeline to ingest data from an Amazon RDS for PostgreSQL database into Amazon S3 using AWS Database Migration Service (AWS DMS). The source database has a high volume of transactions and the engineer needs to capture ongoing changes with minimal impact on the source. The target S3 bucket must store the data in Parquet format for querying with Amazon Athena. Which two actions should the engineer take to meet these requirements? (Choose two.)
Medium1219A data engineer manages an Amazon S3 data lake that holds sensitive customer transaction logs. Compliance requires that all objects be encrypted at rest with keys that the company rotates every 90 days and fully controls, including the ability to immediately revoke access and audit key usage separately from other AWS accounts. The engineer must choose an encryption method that meets these requirements with minimal operational overhead. Which solution should the engineer implement?
Medium1220A data engineer needs to monitor the number of records processed by an AWS Glue ETL job and send an alert if the count drops below a threshold. Which AWS service should be used to create this custom metric?
Easy1221A data engineer is using Amazon Athena to query data stored in an S3 bucket. The queries are running slowly. Which THREE actions can improve query performance?
Hard1222A data engineer needs to monitor the number of records processed by a Kinesis Data Firehose delivery stream and set an alarm if the count drops below a threshold. Which CloudWatch metric should be used?
Easy1223A data engineer is designing a data lake on Amazon S3. Which feature should be used to manage the lifecycle of objects and move them to cheaper storage classes automatically?
Easy1224Refer to the exhibit. A data engineer needs to connect to the Redshift cluster from an EC2 instance in the same VPC. The engineer can ping the EC2 instance but cannot connect to Redshift using the endpoint address and port 5439. What is the most likely cause?
Medium1225A data engineer is building a pipeline to ingest data from an on-premises Oracle database into Amazon S3. The pipeline must capture change data (CDC) in near real-time and handle schema changes. Which TWO AWS services should the engineer use?
Hard1226A company stores raw event files in an Amazon S3 bucket that receives thousands of small objects per hour. An AWS Glue job reads the prefix and writes a compacted Parquet dataset to a curated bucket. Operations reports that the Glue job's runtime keeps growing even though the hourly data volume is constant. Which change is MOST likely to reduce runtime?
Easy1227A data engineer is using AWS Glue to transform data from Amazon S3. The source data is in CSV format with inconsistent date formats across files (e.g., 'MM/DD/YYYY' and 'YYYY-MM-DD'). The engineer needs to standardize all dates to 'YYYY-MM-DD' format in the output. Which AWS Glue transform should the engineer use to achieve this?
Medium1228A data engineer is designing a solution to securely store and rotate database credentials used by an application. The credentials should be automatically rotated every 90 days. Which AWS service should be used?
Hard1229A data engineer is troubleshooting an AWS Glue job that writes data to an S3 bucket. The IAM role attached to the Glue job has the policy shown in the exhibit. The job fails when writing to the 'secrets/' prefix but succeeds when writing to other prefixes. What is the reason for the failure?
Medium1230A data engineer is monitoring an Amazon RDS for PostgreSQL instance. The engineer wants to set up alerts for high CPU utilization and low free storage space. Which AWS services can be used together to achieve this? (Choose TWO.)
Easy1231A data engineer is using AWS Glue to transform data stored in Amazon S3. The security team requires that data in transit between AWS Glue and Amazon S3 be encrypted. The engineer wants to ensure that all connections use TLS. Which action should the engineer take to enforce encryption in transit for AWS Glue jobs accessing S3?
Easy1232A data engineer is using Amazon Kinesis Data Streams to ingest clickstream data. The stream has 10 shards and each record is 50 KB. The engineer notices that the PutRecords API is frequently returning ProvisionedThroughputExceededException errors, even though the total incoming data rate is below the stream's overall capacity. What is the MOST likely cause?
Medium1233A data engineer is designing a real-time analytics pipeline that ingests clickstream data into Amazon Kinesis Data Streams. The data must be stored in Amazon S3 for later analysis with Amazon Athena. The engineer needs the data to be queryable with minimal latency and wants to avoid managing complex ETL jobs. Which solution should the engineer use?
Hard1234A data engineer is running a Spark job on Amazon EMR. The job reads from S3, processes data, and writes to S3. The job is taking longer than expected. The engineer notices that the job is spending a lot of time in the 'GC' (garbage collection) phase. Which configuration change is most likely to improve performance?
Medium1235A company has CSV files in an S3 bucket that need to be converted to Parquet and loaded into a Redshift table daily. The transformation is a simple schema mapping without joins. Which AWS Glue feature is BEST suited for this task?
Easy1236A data engineer is building a data lake on Amazon S3. The engineer needs to store structured data that will be queried by Amazon Athena. The data is currently in CSV format and is partitioned by date. The engineer wants to improve query performance and reduce the amount of data scanned. Which action should the engineer take?
Easy1237A company uses AWS Glue to process data stored in Amazon S3. The security team mandates that all data in transit between AWS Glue and Amazon S3 must be encrypted with TLS. The Glue job connects to S3 using the AWS SDK. Which configuration should the data engineer implement to enforce TLS encryption for the Glue job's S3 connections?
Medium1238A company uses AWS Glue to process data in Amazon S3. The Glue job fails with an error indicating that the partition keys in the catalog do not match the actual S3 partition structure. What is the most likely cause?
Medium1239A data engineer needs to store semi-structured JSON log files from multiple sources and query them using SQL. The data is rarely updated and access frequency is low. Which storage solution is MOST cost-effective?
Easy1240A data engineer is designing a data lake on Amazon S3. The data includes sensitive personally identifiable information (PII). Which combination of services would provide the most comprehensive data protection?
Easy1241A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. The engineer notices that one of the Glue jobs occasionally fails due to a transient network error. The engineer wants to implement a retry mechanism that retries the failed job up to three times with an exponential backoff. Which Step Functions state configuration should the engineer use?
Hard1242A company is using AWS DMS to replicate data from an on-premises Oracle database to Amazon RDS for MySQL. The replication is working, but the target table has a different schema. Which DMS feature should be used to transform the source schema to match the target?
Hard1243A data engineer must load a 500 MB CSV file from Amazon S3 into an existing Amazon Redshift table once per day. The file has a header row and uses a pipe delimiter. The engineer wants the fastest load and the least operational overhead. Which approach should the engineer use?
Easy1244A company stores raw clickstream data in an Amazon S3 bucket and uses AWS Glue crawlers to populate the AWS Glue Data Catalog. Analysts report that new partitions are not appearing in Amazon Athena queries even though new date-based folders exist in S3. The crawler runs successfully each night. What is the MOST likely cause?
Easy1245A company is migrating its on-premises MySQL database to Amazon RDS for MySQL. They want to minimize downtime and ensure data consistency. Which AWS service should be used for the migration?
Easy1246A data engineer is monitoring an Amazon Redshift cluster and notices that queries are taking longer than expected. The engineer checks the system tables and sees that many queries are waiting for 'WLM' resources. What is the most likely cause and recommended fix?
Hard1247A company is using AWS Glue to run ETL jobs that transform data from Amazon S3 to Amazon Redshift. The jobs are failing intermittently with 'Out of Memory' errors. The team wants to resolve this issue without increasing costs significantly. Which TWO actions should the team take?
Medium1248A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources and must be queryable using Amazon Athena. The engineer needs to optimize query performance and reduce costs. Which THREE actions would achieve this?
Hard1249A company uses Amazon EMR to process large datasets stored in Amazon S3. The data is in Parquet format and partitioned by date. The EMR cluster uses Spark SQL for transformations. Recently, the job has been slow and some tasks are failing due to 'java.lang.OutOfMemoryError'. The cluster has 10 core nodes of type m5.xlarge. Which configuration change would MOST improve performance and stability?
Hard1250A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer notices that many queries are waiting in the queue, and the WLM query queue wait time is high. The cluster uses automatic WLM. Which action should the engineer take to improve query throughput?
Medium1251A data engineer manages an AWS Lake Formation governed data lake. Analysts in the finance department must query only the rows in a shared Amazon S3 table where the region column equals 'EMEA', while analysts in the marketing department must see all rows but must not see the customer_email column. Which TWO Lake Formation configurations should the data engineer implement to meet these requirements? (Choose two.)
Medium1252A data engineer is troubleshooting an Amazon Redshift cluster where nightly COPY loads from Amazon S3 are intermittently slow and sometimes fail with 'S3ServiceException' errors. The engineer suspects the cluster's network configuration and load design are contributing. Which TWO actions should the engineer take to improve load performance and reliability? (Choose two.)
Hard1253A company uses Amazon Redshift for data warehousing. They notice that queries are running slowly, and the STL_LOAD_ERRORS table shows many 'Parse error' entries. The data is loaded from Amazon S3 using COPY commands. What is the MOST likely cause of the parse errors?
Hard1254Which TWO actions can help improve the read performance of an Amazon DynamoDB table that is experiencing throttling? (Choose two.)
Medium1255A company uses Amazon EMR to run Spark jobs on a transient cluster. The jobs process data from S3 and write results back to S3. The team wants to reduce costs by optimizing the cluster. Which action should the team take?
Easy1256A data engineer runs the above SQL commands on an Amazon Redshift cluster. The table 'users' is created with DISTSTYLE EVEN. What is the effect of the DISTSTYLE EVEN on query performance?
Easy1257A data engineering team uses AWS Glue ETL jobs to process data daily. They notice that job run times are increasing as data volume grows. Which action will most effectively improve performance without changing the code?
Medium1258A company uses AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by date. The company wants to improve query performance in Amazon Athena and reduce costs. Which THREE actions should the company take? (Choose THREE.)
Medium1259A company is designing a data pipeline using Amazon Kinesis Data Streams. The data includes personally identifiable information (PII). The security team requires that data be encrypted at rest using a customer-managed KMS key. How should the data engineer configure the Kinesis stream?
Hard1260A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket with millions of small JSON files. The job is running slowly and often fails with out-of-memory errors. The engineer needs to improve performance and reliability. What should the engineer do?
Hard1261A company is using AWS Lake Formation to manage access to a data lake in S3. They want to grant a data analyst access to specific columns in a table, but not to the entire table. Which Lake Formation feature should be used?
Medium1262A data engineer is designing an AWS Glue ETL job that reads data from an Amazon S3 bucket containing many small JSON files and writes the output to Amazon S3 in Parquet format. The engineer wants to improve read performance and reduce the number of output files. Which two actions should the engineer take? (Choose two.)
Hard1263A data engineer needs to grant a data scientist access to query a Glue Data Catalog database but must prevent the data scientist from seeing the underlying S3 data locations. Which approach should be used?
Hard1264A company uses AWS Glue to run ETL jobs daily. The data is stored in S3 as Parquet files partitioned by date. Recently, jobs have failed with the error 'No such file or directory' for certain partitions. What is the MOST likely cause?
Easy1265A data engineer is setting up a data pipeline using AWS DMS to migrate data from an on-premises database to Amazon RDS for MySQL. The data must be encrypted in transit. Which TWO options can the engineer use? (Choose TWO.)
Easy1266A data engineer is designing a data ingestion pipeline for real-time clickstream data from a website. The data must be ingested with low latency (seconds) and made available for multiple consumer applications, including a dashboard that refreshes every minute and a machine learning model that processes data in near-real-time. The engineer needs to choose a streaming ingestion service. Which TWO services meet these requirements? (Select TWO.)
Medium1267A company uses Amazon Kinesis Data Streams to ingest real-time logs from thousands of applications. The data must be transformed and enriched with reference data from Amazon S3 before being stored in Amazon S3 in Parquet format. The transformation logic is stateful and requires exactly-once processing. Which AWS service should the data engineer use to perform the transformation?
Medium1268A company wants to enforce that all data in an S3 bucket is encrypted at rest using AWS KMS. Which bucket policy condition key should be used?
Easy1269Match each AWS monitoring tool to its primary use.
Medium1270A company uses Amazon S3 as a data lake. A data engineer needs to ensure that all objects uploaded to the 'incoming' prefix are automatically encrypted at rest using AWS KMS with a specific customer managed key. What is the simplest way to enforce this?
Easy1271A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an Amazon RDS for PostgreSQL database to an Amazon S3 bucket. The source table has a primary key and the engineer needs near-real-time change data capture (CDC) with minimal impact on the source. Which DMS task setting should the engineer configure to meet these requirements?
Hard1272Arrange the steps to implement data encryption at rest for an Amazon Redshift cluster using AWS KMS.
Medium1273A company uses Amazon Redshift for data warehousing. The security team requires that all data loading into Redshift be encrypted in transit. Which configuration ensures this requirement is met?
Medium1274A data engineer stores Apache Parquet files in an Amazon S3 data lake partitioned by dt=YYYY-MM-DD. Analysts query the data with Amazon Athena, and monthly reports that scan one month of data are slow and expensive. The engineer confirms that queries filter on the dt column. Which action will MOST effectively reduce the amount of data scanned by these reports?
Medium1275A data engineer needs to ingest data from an Amazon S3 bucket into Amazon Redshift for analytics. The data is in CSV format and the Redshift table already exists. Which service can be used to perform this ingestion with minimal configuration?
Easy1276A data engineer is troubleshooting an issue where an Amazon Redshift query returns an error: 'ERROR: permission denied for relation table_name'. The user has been granted SELECT on the table. What is the most likely cause?
Hard1277A data engineer is using AWS Lake Formation to manage access to a data lake in Amazon S3. The engineer needs to grant a specific IAM role access to only the columns containing non-sensitive data in a table stored in the AWS Glue Data Catalog. The role should not have access to sensitive columns. What should the engineer do?
Medium1278A logistics company stores shipment tracking events in an Amazon DynamoDB table. The table uses a partition key of shipment_id and a sort key of event_timestamp. Analysts frequently run queries that filter by shipment_id and a range of event_timestamp values. The data engineer must ensure these queries are efficient and consume minimal read capacity. What should the data engineer do?
Easy1279A data engineer needs to store time-series data from IoT devices. The data is write-heavy and requires low-latency queries by device ID and timestamp. The data volume is expected to grow to terabytes. Which AWS database service is most suitable?
Easy1280A company has an Amazon RDS for MySQL database that is experiencing performance issues due to a large number of read requests. The application is read-heavy and can tolerate eventually consistent reads. Which action will reduce the load on the primary database with the least operational overhead?
Hard1281A data engineer needs to ensure that data in an Amazon S3 bucket is not publicly accessible. Which TWO measures should the engineer implement? (Choose TWO.)
Medium1282A data engineer must give an AWS Glue ETL job access to an S3 bucket that is encrypted with SSE-KMS using a customer managed key. The Glue job runs under an IAM role. The security team wants the least-privilege permissions required for the job to read and write objects in that bucket. Which TWO actions must be included in the IAM role's policy? (Choose two.)
Medium1283A data engineer is designing a data pipeline that ingests data from an on-premises database into Amazon S3 using AWS Database Migration Service (DMS). The data must be encrypted at rest in S3 using SSE-S3. The engineer also needs to track changes to the source database in real time. Which DMS configuration should the engineer use?
Easy1284A company is running an Amazon EMR cluster with Spark for data processing. The data engineer wants to automatically scale the core and task nodes based on the YARN memory and CPU utilization. Which scaling metric should the engineer use for the EMR managed scaling policy?
Medium1285A data engineer is monitoring an Amazon Kinesis Data Stream used to ingest clickstream data. The engineer notices that the stream's 'WriteProvisionedThroughputExceeded' metric is frequently above zero. Which TWO actions could help mitigate this issue? (Choose TWO.)
Easy1286A data engineer is managing an Amazon S3 data lake that contains millions of small files. The engineer needs to optimize query performance in Amazon Athena and reduce costs. The data is stored in Parquet format and is partitioned by date. Which action should the engineer take to improve performance and reduce costs?
Hard1287A company is migrating a legacy data warehouse to Amazon Redshift. They need to choose a distribution style to minimize data movement during joins. Which THREE factors should they consider?
Hard1288A data engineer is building a data pipeline that ingests data from Amazon S3 into Amazon Redshift. The data is in CSV format and includes a timestamp column. The pipeline should load only new data incrementally. Which approach is most efficient?
Medium1289A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is consumed by a custom consumer application that writes to Amazon S3 every 5 minutes. The consumer is falling behind and processing lag is increasing. Which action is MOST effective to reduce the lag?
Easy1290A data engineer is configuring an AWS Glue ETL job that reads semi-structured JSON event logs from Amazon S3 and must flatten nested arrays into relational columns before writing to Amazon Redshift. The job must run reliably without writing custom serialization code. Which approach should the data engineer take?
Medium1291A company uses Amazon Redshift for its data warehouse. The data engineering team loads data daily from Amazon S3 using COPY commands. Recently, the load performance has degraded because the S3 bucket contains many small files. The team needs to optimize the COPY operation to improve performance. Which approach should they take?
Easy1292A data engineer is troubleshooting a failed AWS Glue Crawler. The crawler logs show 'Insufficient permissions to access S3 bucket'. What should the engineer do to resolve this?
Easy1293A company has a nightly batch job that processes 100 GB of data from an Amazon S3 bucket and loads it into an Amazon Redshift table. The job currently runs on an Amazon EMR cluster. Which service would reduce operational overhead while providing similar functionality?
Easy1294A data engineer is running an Amazon Athena query that scans a large amount of data in Amazon S3, resulting in high costs. The data is stored in Parquet format in a partitioned table. Which strategy would be MOST effective in reducing the amount of data scanned?
Medium1295A data engineer is designing a DynamoDB table for an application that requires strongly consistent reads and supports a global secondary index (GSI). The engineer needs to ensure that queries on the GSI return the most up-to-date data. Which statement about DynamoDB read consistency is correct?
Hard1296A company is running a Redshift cluster and wants to improve query performance for a frequently used dashboard. Which THREE approaches are recommended?
Hard1297Which TWO AWS services can be used to ingest streaming data into Amazon S3? (Choose two.)
Easy1298A data engineer is designing a real-time streaming pipeline to ingest clickstream data from a website into Amazon S3. The data must be transformed before storage. Which TWO AWS services can be used together to build this pipeline? (Choose TWO.)
Easy1299A data engineer is using AWS Glue to read a large dataset from Amazon S3 and write it to Amazon Redshift. The job intermittently fails with 'Communication link failure' errors during the write phase. The dataset is several hundred gigabytes and the Redshift cluster is under heavy query load. Which change is MOST likely to resolve the failures while preserving data integrity?
Hard1300A company uses AWS Glue ETL jobs to transform data from Amazon S3 to Amazon Redshift. The job reads JSON files, applies schema mapping, and writes to a Redshift table. Recently, the job started failing with memory errors. The data volume has increased tenfold. Which approach should a data engineer take to resolve this issue with minimal code changes?
Medium1301A data engineer is tasked with setting up a data pipeline that moves data from an on-premises Oracle database to Amazon S3 every hour. The network bandwidth is limited, and the engineer needs to ensure data consistency. Which AWS service should the engineer use?
Easy1302A company needs to enforce encryption at rest for all data stored in Amazon S3. Which of the following are valid methods to achieve this? (Choose TWO.)
Medium1303Refer to the exhibit. A data engineer applies the following S3 bucket policy to an S3 bucket. What does this policy enforce?
Medium1304A data engineer needs to store encryption keys used for protecting data in Amazon S3 and automatically rotate them every year. Which service should be used?
Easy1305A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Redshift. The migration uses a full load plus change data capture (CDC). During the CDC phase, the engineer notices that some updates are not being applied to the target Redshift tables. The DMS task logs show no errors. What is the MOST likely cause?
Hard1306A data engineer is designing a data lake on Amazon S3 that will store sensitive financial data. The engineer needs to implement encryption at rest and ensure that only authorized users can access the data. Which TWO actions should the engineer take to meet these requirements? (Choose TWO.)
Medium1307A data engineer is building a data ingestion pipeline using AWS Glue. The source is an Amazon DynamoDB table, and the target is an Amazon S3 data lake in Parquet format. The pipeline must handle large volumes and ensure exactly-once processing. Which THREE features should the engineer use together to achieve this? (Choose THREE.)
Hard1308A company wants to enforce that all data in Amazon S3 is encrypted at rest. They want to automatically reject any PUT request that does not include encryption headers. What S3 feature should they use?
Easy1309A data engineer uses AWS Glue to process data from S3. The Glue job frequently fails with 'Out of Memory' errors. The job reads several large compressed files. What is the MOST effective way to resolve this issue without changing the code?
Medium1310A company is ingesting streaming data from thousands of IoT devices into Amazon Kinesis Data Streams. The data is processed by a Kinesis Data Analytics application. Recently, the application started reporting high iterator age (millisBehindLatest). Which action would BEST reduce the iterator age?
Medium1311A company stores sensitive customer data in an S3 bucket. The data engineer needs to ensure that all data is encrypted at rest. Which S3 feature should be enabled?
Easy1312A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data must be transformed in real-time and then stored in Amazon S3. Which AWS service should be used to perform the transformation?
Easy1313A company uses AWS DMS to migrate an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. After initial load, ongoing replication is set up. The replication task shows 'Task status: failed with error: The specified LSN is not available in the source database logs.' What is the most likely cause?
Medium1314A data pipeline uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The delivery stream is configured with a buffer size of 5 MB and a buffer interval of 60 seconds. The team notices that the S3 objects are much smaller than 5 MB. What is the most likely explanation?
Hard1315Refer to the exhibit. A data engineer runs the above CLI command to find files smaller than 1000 bytes in a bucket. The command returns an empty array, but the engineer knows there are small files. What is the issue?
Easy1316A data engineer is troubleshooting a slow Amazon Redshift query that joins several large tables. The query plan shows a large number of broadcasts. Which design change would most likely reduce the broadcast operations?
Medium1317A data engineer is using AWS Glue to process a large dataset stored in Amazon S3. The dataset is partitioned by year, month, and day. The engineer notices that the Glue job is taking a long time and consuming many DPUs. The job reads all partitions, filters the data, and writes the result to another S3 location. The engineer wants to optimize the job to process only the required partitions and reduce cost. Which action should the engineer take?
Hard1318A data engineer needs to share a dataset from an S3 bucket in Account A with users in Account B. The dataset must remain encrypted at rest with an S3-managed key. What is the MOST secure way to grant cross-account access?
Easy1319A company uses Amazon Kinesis Data Analytics for real-time anomaly detection on clickstream data. The application uses a sliding window of 1 minute. The data engineer notices that the application is producing incorrect results because late-arriving records are not being handled properly. What should the data engineer do to ensure late records are included in the window calculations?
Hard1320A data engineer is using AWS Glue to read data from an Amazon Kinesis Data Stream. The Glue job is configured to process the stream in micro-batches. The engineer notices that the job is not processing all records and sometimes skips data. The Kinesis stream has multiple shards, and the Glue job is using the 'kinesis' connection type. What is the most likely cause of the missing records?
Hard1321A data engineer is orchestrating a multi-step ingestion workflow where CSV files land in Amazon S3, an AWS Glue job transforms them, and the output is loaded into Amazon Redshift. The engineer wants conditional branching, retry logic, and the ability to pass parameters between steps. Which AWS service should be used to orchestrate this workflow?
HardOther domains
All DEA-C01 exam domains
Frequently asked questions
- What does the troubleshooting domain cover on the DEA-C01 exam?
- troubleshooting questions test whether you can apply the concept in context, not just recognise a definition.
- How many questions are in this domain?
- This page lists all 1321 troubleshooting questions in the DEA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only troubleshooting questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.