Candidates must configure S3 lifecycle policies, DynamoDB capacity and GSIs, Glue Data Catalog partitions, and Lake Formation grants. Get S3 storage class transitions and Lake Formation column-level permissions right, since most scenario questions hinge on least-privilege access and cost-optimal storage.
Start practicing
Data Store Management — choose a session length
Free · No account required
Domain overview
Data Store Management is 26% of DEA-C01 and covers choosing, configuring and securing AWS storage for analytics workloads. Expect scenario questions on S3 storage classes and lifecycle rules, DynamoDB capacity and indexes, Lake Formation permissions, Glue Data Catalog tables and partitions, and KMS encryption choices across Redshift, RDS and S3.
Exam objectives
Selecting S3 storage classes, lifecycle transitions and Intelligent-Tiering for cost and access patterns
Configuring DynamoDB partition keys, GSIs/LSIs, on-demand versus provisioned capacity and DynamoDB Streams
Managing Glue Data Catalog databases, tables, partitions and crawler classification of S3 data
Applying Lake Formation grants, S3 bucket policies and KMS keys for fine-grained data access control
Choosing S3 Glacier Deep Archive for data needing millisecond retrieval, ignoring the hours-long restore time and retrieval charges
Using a low-cardinality DynamoDB partition key, causing hot partitions and throttling instead of even read/write distribution
Granting IAM permissions on S3 but forgetting Lake Formation table grants, so Athena and Redshift Spectrum queries still fail
Click any question to see the full explanation and answer options, or start a focused practice session above.
A company has an Amazon RDS for MySQL DB instance with read replicas. The primary DB instance fails. What is the correct procedure to promote a read replica to become the new primary?
2A company uses Amazon DynamoDB for a gaming application. They need to store player session data that expires after 24 hours. Which DynamoDB feature should they use to automatically delete expired items?
3An e-commerce application uses Amazon ElastiCache for Redis to cache product catalog data. The cache currently uses lazy loading. The team wants to ensure that frequently accessed product data is always fresh. Which caching strategy should they implement?
4Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket. What is the effect of this policy?
5Refer to the exhibit. A data engineer runs the above AWS CLI command to view the table metadata in the AWS Glue Data Catalog. The data is stored as CSV in S3 with partitions by year and month. When querying the table using Amazon Athena, no data is returned. What is the most likely cause?
6Which THREE storage classes in Amazon S3 are designed for infrequently accessed data with millisecond retrieval times? (Select THREE.)
7A company stores time-series sensor data in Amazon S3. They need to query the data using SQL with minimal latency and no infrastructure management. Which service should they use?
8A data engineer needs to migrate an on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which approach should they use?
9A company is designing a data lake on Amazon S3. Which TWO strategies improve query performance for Amazon Athena?
10A company is migrating a legacy data warehouse to Amazon Redshift. They need to choose a distribution style to minimize data movement during joins. Which THREE factors should they consider?
11A company runs a multi-AZ Amazon RDS for PostgreSQL instance. They need to run a one-time analytical query that will take several hours and consume significant I/O. The query should not impact the primary workload. What should the data engineer do?
12A company is migrating its on-premises MySQL database to Amazon RDS for MySQL. They want to minimize downtime and ensure data consistency. Which AWS service should be used for the migration?
13A data engineer is troubleshooting a slow-running query on an Amazon Redshift cluster. The query involves joining two large tables. The engineer notices that the query plan shows a large number of distribution and broadcast operations. Which design change would most likely improve query performance?
14A company runs a real-time analytics platform using Amazon Kinesis Data Streams with a shard count of 10. The data is consumed by an AWS Lambda function that writes to an Amazon DynamoDB table. The DynamoDB table has a partition key of 'user_id' and a sort key of 'timestamp'. The table is provisioned with 5000 RCUs and 5000 WCUs. Recently, the application experienced increased write latency and throttling errors (ProvisionedThroughputExceededException) on the DynamoDB table. The CloudWatch metrics show that ConsumedWriteCapacityUnits averages 4500 with occasional spikes to 6000. The Lambda function’s concurrency is set to 1000. The data engineer suspects the issue is due to hot partitions. Upon investigation, the engineer finds that a small number of users generate a disproportionately large amount of data. Which course of action would best resolve the throttling while minimizing cost?
15Order the steps to set up an Amazon EMR cluster for processing data in S3 using Spark.
16Match each AWS Glue component to its role.
17Match each AWS monitoring tool to its primary use.
18A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources and needs to be partitioned by year, month, day, and event type for efficient querying with Amazon Athena. Which S3 key prefix structure is most appropriate?
19A data engineer needs to store semi-structured JSON data from IoT devices. The data is written once, read rarely, but must be queryable using SQL. The storage cost must be minimized. Which storage solution should the engineer choose?
20A company uses Amazon DynamoDB as the primary data store for a web application. The application experiences occasional throttling on write requests. The data engineer needs to implement a solution that handles throttling gracefully without losing data. Which approach should the engineer use?
21A financial services company stores transaction data in Amazon RDS for PostgreSQL. The company requires that all changes to the database be logged for audit purposes, including before and after images of updated rows. Which feature should the data engineer enable?
22A data engineer is setting up Amazon S3 bucket policies for a data lake. Which TWO statements are true regarding S3 bucket policies? (Choose TWO.)
23A data engineer applies the above IAM policy to a user. The user attempts to upload an object to the bucket 'my-data-lake' without specifying server-side encryption. What will happen?
24A data engineer runs the above CLI command to describe the DynamoDB table 'Orders'. The table has a partition key 'OrderID' and sort key 'CustomerID'. Which query operation is most efficient for retrieving all orders for a specific customer?
25A data engineer is troubleshooting an access denied error when an AWS Lambda function tries to decrypt an object encrypted with the KMS key 'abc123'. The Lambda function's execution role has the above policy attached. What is the likely cause of the error?
26A company wants to store data from thousands of IoT devices with varying data rates. The data must be stored in a schema-on-read fashion and support SQL queries. Which AWS service should be used?
27A company runs a data warehouse on Amazon Redshift. Queries are slow, and the team suspects data distribution is skewed. Which approach would best help identify distribution skew?
28A company needs to store JSON documents that are frequently read and written by a web application. The data must be highly available and durable across multiple Availability Zones. Which AWS database service meets these requirements?
29A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The migration must have minimal downtime. Which service should be used to perform the migration?
30A company uses Amazon Redshift for analytics. The data engineering team wants to improve query performance for frequently used aggregate queries. Which TWO actions would help achieve this?
31A company stores sensitive data in Amazon S3. The security team requires encryption at rest and that the encryption keys are managed by the company using AWS KMS. The data is frequently accessed by multiple AWS services. Which THREE steps should be taken to meet these requirements?
32A company is designing a data lake on Amazon S3. The data includes CSV files, Parquet files, and images. The data engineering team needs to catalog the metadata and enable SQL queries. Which TWO AWS services should be used together?
33A company uses Amazon DynamoDB for a gaming application. The application experiences throttling during peak hours. The table's read and write capacity is provisioned. Which TWO actions can reduce throttling?
34A company is migrating a large Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime and preserve data consistency. Which THREE AWS services or features should be used?
35A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data is immutable and must be stored cost-effectively. Which AWS service should the engineer use?
36A company runs an Apache Spark job on Amazon EMR that writes output to an S3 bucket. The job fails with the error 'S3AccessDeniedException' when writing the final output, but earlier stages succeed. The EMR cluster uses a service role and an instance profile. The S3 bucket policy allows access from the VPC only. What is the MOST likely cause?
37A data engineer is designing a data lake on Amazon S3. The data includes customer PII that must be encrypted at rest. The company also requires that the encryption keys be rotated automatically every year. Which encryption solution should the engineer use?
38A company needs to store application log files for 90 days for compliance. The logs are generated continuously and are rarely accessed after 30 days. The data engineer must minimize storage costs. Which storage solution should the engineer choose?
39A company is designing a data store for IoT sensor data that is written once and never updated. The data must be stored with high durability and low cost. Which TWO AWS storage services are most suitable? (Choose TWO.)
40A data engineer applies the bucket policy shown in the exhibit to an S3 bucket. The bucket contains sensitive data that must be encrypted at rest and accessed only over HTTPS. Which of the following statements is true?
41A company has an Amazon DynamoDB table with a provisioned write capacity of 1000 WCU. During a flash sale, the write traffic spikes to 5000 WCU for 10 minutes. The table is not auto-scaled. Which action should the data engineer take to handle the spike without throttling?
42A data engineer is designing a data lake on Amazon S3 and needs to ensure that objects are automatically encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement achieves this?
43A company uses Amazon Redshift for its data warehouse. The data engineer notices that queries are slow on a large table that is frequently filtered on a column 'transaction_date'. Which optimization technique best improves query performance?
44A data engineer is migrating an on-premises MongoDB database to Amazon DocumentDB. Which migration strategy minimizes downtime?
45A company needs to store files that are accessed by multiple EC2 instances in a VPC. The files must be concurrently accessible and durable. Which storage solution should the data engineer choose?
46Which TWO actions are recommended for securing data at rest in Amazon S3? (Choose two.)
47A data engineer needs to store JSON documents that are accessed by a serverless application using AWS Lambda. The documents are frequently updated and need low latency (single-digit milliseconds) for read and write operations. Which AWS service should the engineer use?
48A data engineer needs to store semi-structured JSON logs from multiple sources in a centralized data store for querying using SQL. The logs are immutable and need to be retained for 90 days. Which AWS service should be used?
49A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow due to many small inserts. Which technique would improve write performance?
50A company uses Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed key that is rotated annually. Which encryption option should be used?
51A data engineer is troubleshooting a slow-running query on Amazon Redshift. The query scans a large table but returns few rows. Which diagnostic step should be taken first?
52A company uses DynamoDB with global tables in two AWS Regions. The data engineer observes that a write to the table in us-east-1 is not immediately visible in a read from eu-west-1. What is the most likely reason?
53A data engineer needs to store event data from IoT devices that arrives in bursts. The data is key-value and requires single-digit millisecond read and write latency. The engineer also needs to run complex analytical queries on the data for reporting. Which TWO services should be used together? (Choose TWO.)
54Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket named data-lake-bucket. The engineer wants to allow only GET requests from the corporate network (10.0.0.0/16) over HTTPS. However, users report that they cannot access objects even when connected to the corporate network. What is the issue?
55Refer to the exhibit. A data engineer needs to connect to the Redshift cluster from an EC2 instance in the same VPC. The engineer can ping the EC2 instance but cannot connect to Redshift using the endpoint address and port 5439. What is the most likely cause?
56A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for ad-hoc querying using SQL. Which AWS service should be used?
57A company uses Amazon S3 as its data lake. A data engineer needs to enforce encryption of data at rest using server-side encryption with AWS KMS. Which S3 bucket property should be configured?
58A data engineer notices that an Amazon Redshift cluster’s storage usage is increasing rapidly due to many UPDATE and DELETE operations. The engineer needs to reclaim storage space and improve query performance. Which action should be taken?
59A data engineer is troubleshooting an Amazon Redshift cluster that is running out of disk space. The engineer runs STV_PARTITIONS and notices that some slices have significantly more data than others. What is the most likely cause and solution?
60A data engineer needs to store transaction data that requires strong consistency, ACID transactions, and complex join queries. Which AWS service is most appropriate?
61A data engineer is designing a multi-Region disaster recovery solution for an Amazon DynamoDB table. The table must be available in a secondary Region with minimal data loss and automatic failover. Which feature should be used?
62Which THREE factors should be considered when choosing a partition key for an Amazon DynamoDB table?
63A data engineer needs to store semi-structured JSON transaction logs for analytics. The logs are written once and rarely accessed. The storage must be cost-effective. Which AWS service should be used?
64A company uses Amazon Redshift for a data warehouse. They notice that queries are slow due to heavy data skew. Which optimization technique should be applied first?
65A company has a DynamoDB table with a partition key of 'user_id' and a sort key of 'timestamp'. They need to query all items for a user within a date range. Which query operation should be used?
66A data engineer needs to store streaming data from IoT devices for real-time analytics. The data has a fixed schema and requires low-latency queries. Which AWS service should be used?
67A data engineer is designing a multi-region disaster recovery solution for Amazon RDS for PostgreSQL. The primary region must have a standby in a different Availability Zone, and the secondary region must have a readable replica that can be promoted in case of failure. Which configuration meets these requirements?
68A company stores log files in Amazon S3. They want to automatically move logs older than 90 days to S3 Glacier Deep Archive to reduce costs. Which S3 feature should be used?
69Which TWO actions can help optimize Amazon S3 storage costs for a data lake? (Choose two.)
70Which TWO statements about Amazon Redshift data distribution are correct? (Choose two.)
71Refer to the exhibit. A data engineer ran the CLI command to check the configuration of an RDS instance named 'mydb'. Which statement accurately describes the current configuration?
72A company is using Amazon S3 to store critical data and needs to ensure that objects are automatically transitioned to S3 Glacier Deep Archive after 180 days to reduce costs. Which S3 lifecycle action should be configured?
73A company is using Amazon DynamoDB for a gaming application. They want to store player session data that expires after 24 hours. Which DynamoDB feature should be used?
74A company has an Amazon Redshift cluster that stores petabytes of data. Queries are experiencing high disk usage due to large intermediate results. The data engineer needs to improve query performance without adding more nodes. Which action should the engineer take?
75A company is using Amazon S3 to store sensitive data. They need to ensure that all objects are encrypted at rest. Which combination of actions should be taken? (Choose TWO.)
76A company uses Amazon Redshift for analytics. They notice that some queries are slow due to data redistribution. The data engineer wants to minimize data movement across nodes. Which table design strategy should be used? (Choose TWO.)
77A company is using Amazon S3 for data lake storage. They need to query the data directly using SQL without loading it into a database. Which AWS service should be used?
78A data engineer sees this AWS Glue table definition in the Data Catalog. The engineer wants to query this table with Amazon Athena, but the query returns zero rows. What is the MOST likely cause?
79A CloudFormation template includes this IAM policy for a cross-account S3 upload use case. What is the purpose of the condition?
80A company is using Amazon RDS for MySQL and needs to automate backups with a retention period of 35 days. They also want to be able to restore to any point within the retention period. Which configuration should be used?
81A company is storing sensitive user data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using a customer-managed key stored in AWS KMS. The bucket policy must deny any PUT request that does not include the appropriate encryption header. Which bucket policy condition key should be used?
82A data engineer needs to set up a new Amazon RDS for MySQL database for a web application. The application experiences variable read traffic and requires low read latency. The engineer needs to minimize downtime during maintenance and provide read scalability. Which configuration meets these requirements?
83An e-commerce company uses Amazon DynamoDB as the primary data store for its product catalog. The table has a simple primary key (ProductID) and handles 10,000 writes per second during peak hours. Recently, the engineering team noticed increased write latency and throttled requests during peak times. The table's provisioned write capacity is set to 12,000 WCU. What is the most likely cause of the throttling?
84A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for later analysis using Amazon Athena. The logs are generated continuously, and the total volume is about 1 TB per day. The data must be queryable within minutes of arrival. Which storage solution is most appropriate?
85A data engineer created the IAM policy shown in the exhibit. The engineer then attempts to upload an object to 'my-bucket' using the AWS CLI with the command: aws s3 cp file.txt s3://my-bucket/ --sse aws:kms. The upload fails with an 'AccessDenied' error. What is the most likely cause?
86A data engineer is configuring an Amazon S3 lifecycle policy to transition objects to S3 Glacier Deep Archive after 90 days. The bucket receives new objects daily. The engineer wants to ensure that objects are not deleted before 90 days. Which lifecycle action should be used?
87A data engineer is troubleshooting slow query performance on an Amazon Redshift cluster. The cluster has 10 nodes and is using automatic distribution style. The engineer suspects that data distribution is causing excessive data movement. Which steps should the engineer take to diagnose and resolve the issue? (Choose THREE.)
88A data engineer reviewed the S3 lifecycle policy shown in the exhibit. The engineer notices that objects under the 'logs/' prefix are being deleted after 365 days. The business requirement is to retain logs for at least 5 years. What should the engineer change in the lifecycle policy?
89A data engineer is troubleshooting a slow Amazon Redshift query that joins several large tables. The query plan shows a large number of broadcasts. Which design change would most likely reduce the broadcast operations?
90A company uses Amazon DynamoDB with provisioned capacity. During a sales event, write traffic spikes and some requests receive ProvisionedThroughputExceeded exceptions. The reads are within limits. The data engineer needs to minimize latency for the spike without manual intervention. Which solution is MOST cost-effective?
91A data engineer needs to store semi-structured JSON log files from multiple sources and query them using SQL. The data is rarely updated and access frequency is low. Which storage solution is MOST cost-effective?
92A data engineer is designing a data lake on Amazon S3. The data consists of sensitive personally identifiable information (PII) that must be encrypted at rest. The company requires that encryption keys be rotated every 90 days and that access to the keys be logged. Which encryption solution meets these requirements?
93A company is migrating an on-premises Apache Cassandra database to Amazon Keyspaces. The database has a table with a partition key of 'user_id' and a clustering column of 'timestamp'. The application frequently queries the last 10 records for a given user. Which table design in Keyspaces would provide the BEST query performance for this access pattern?
94A company uses Amazon RDS for PostgreSQL. The data engineer needs to ensure that the database is automatically backed up and that backups are retained for 35 days. What is the simplest way to achieve this?
95A data engineer is designing a data lake on Amazon S3 that will be accessed by multiple AWS Glue ETL jobs. The engineer needs to ensure that the data is organized efficiently for querying and that sensitive columns are masked for certain users. Which TWO actions should the engineer take? (Choose TWO.)
96A data engineer notices that an Amazon Redshift cluster is experiencing slow query performance. The engineer suspects that tables are not properly sorted. Which diagnostic query should the engineer run to identify unsorted rows?
97A company is using Amazon S3 as a data lake. The data engineer needs to ensure that all objects uploaded to a specific bucket are automatically replicated to a bucket in another AWS Region for disaster recovery. Which configuration should the engineer implement?
98A data engineer needs to store and analyze time-series data from IoT devices. The data volume is 10 GB per day, and the queries are mostly on the most recent 7 days of data. The engineer wants to minimize storage costs while retaining historical data for 1 year. Which combination of AWS services is most cost-effective?
99A company is using Amazon EMR to process large datasets stored in Amazon S3. The data engineer wants to reduce the time it takes to read data from S3 by optimizing the data format. Which file format should the engineer recommend?
100A company is using Amazon Redshift for data warehousing. The data engineer notices that the STL_ALERT_EVENT_LOG table shows many 'missing statistics' alerts. What is the best course of action to address this issue?
101Which THREE of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose three.)
102Which TWO actions can help improve query performance in Amazon Redshift? (Choose two.)
103A company runs a real-time analytics platform on Amazon ECS that ingests streaming data from Amazon Kinesis Data Streams, processes it, and stores results in Amazon DynamoDB. The data volume spikes unpredictably, causing DynamoDB to throttle write requests. The application uses on-demand capacity mode. The data engineer notices that the throttling occurs on a specific partition due to a hot key. The hot key is a customer ID that receives a disproportionate number of writes. The application cannot change the partition key design immediately. The engineer needs to reduce throttling while maintaining low latency. Which solution is most effective?
104A company uses Amazon DynamoDB as its primary data store for a web application. The application experiences high latency during peak hours. The data engineer notices that the table has a large number of items with the same partition key. Which DynamoDB feature should the engineer use to improve performance?
105A data engineer needs to store semi-structured JSON data that is accessed infrequently but must be retrievable within minutes. The data is generated by IoT devices and each object is about 500 KB. The engineer wants the most cost-effective storage solution. Which AWS service should be used?
106A data engineer needs to store log files from multiple applications in a centralized location. The logs are generated in JSON format and each log entry is about 1 KB. The engineer needs to query the logs occasionally using SQL-like queries. Which AWS service is most appropriate?
107A data engineer is migrating a large Oracle data warehouse to Amazon Redshift. The engineer needs to ensure optimal performance. Which TWO practices should the engineer follow?
108A company is using an Amazon RDS for MySQL database for its e-commerce platform. During a recent flash sale, the database experienced high read traffic, causing slow query performance. The company needs a solution that offloads read traffic with minimal application changes. Which action should be taken?
109A data engineering team is designing a data lake on Amazon S3. They need to store raw data in its original format and transformed data in Parquet. The data is accessed by multiple analytics services, including Amazon Athena and Amazon Redshift Spectrum. Compliance requirements mandate that all data be encrypted at rest with AWS KMS and that the encryption keys be rotated every 90 days. Which S3 bucket configuration meets these requirements?
110A company uses Amazon DynamoDB as the primary data store for a gaming application. The application stores user profiles and game state. During peak hours, the application experiences throttling on writes to the UserProfiles table. The table's read capacity is underutilized. Which solution should resolve the write throttling?
111A company is running a data warehouse on Amazon Redshift. The data engineering team notices that query performance has degraded over time. They suspect that data distribution is causing excessive data movement between nodes. The table is joined frequently on the customer_id column. Which column should be chosen as the distribution key to optimize join performance?
112A company uses Amazon S3 to store sensitive financial data. The security team requires that all objects be encrypted at rest using AWS KMS with a customer-managed key. Additionally, they want to audit all KMS decrypt calls for compliance. Which configuration should be used to meet these requirements?
113A data engineer needs to store time-series data from IoT devices. The data is write-heavy and requires low-latency queries by device ID and timestamp. The data volume is expected to grow to terabytes. Which AWS database service is most suitable?
114A company is migrating its on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is 2 TB and runs on a single server. Which AWS service should be used for the migration?
115A company is using Amazon DynamoDB with auto scaling enabled. During a marketing campaign, write traffic spikes, and some write requests fail with ProvisionedThroughputExceededException. The auto scaling policy has a target utilization of 70% and a maximum capacity that is high enough. What is the most likely cause of the throttling?
116A company is using Amazon Redshift for its data warehouse. The data engineering team needs to improve query performance for a large fact table that is frequently joined with multiple dimension tables. Which THREE strategies should be considered?
117A data engineer runs the AWS CLI command to retrieve the lifecycle configuration of the 'my-data-lake' bucket. The output is shown in the exhibit. What is the effect of this lifecycle policy?
118A company stores its application logs in Amazon S3. The logs are generated daily and need to be retained for 3 years for compliance. The logs are accessed frequently for the first 30 days, occasionally for the next 6 months, and rarely after that. The data engineering team wants to minimize storage costs while ensuring that logs are available for retrieval within 12 hours for the first 6 months and within 48 hours after that. The team also wants to automatically delete logs after 3 years. Which lifecycle policy should the team implement?
119A data engineering team is designing a data lake on Amazon S3 for storing sensor data from IoT devices. The data is written in near real-time and needs to be queried using Amazon Athena. Which TWO configurations should the team implement to optimize query performance and minimize costs?
120A company uses Amazon S3 to store historical stock market data as CSV files. They run daily Amazon Athena queries to generate reports. Recently, the finance team reported that queries are timing out and costs have increased significantly. The data engineering team notices that the S3 bucket contains thousands of small files (average 100 KB) due to a misconfigured ingestion pipeline. They need to improve query performance and reduce costs without changing the existing reporting schedule. The team has access to AWS Glue and can create new tables. Which solution should they implement?
121A company uses Amazon Redshift for its data warehouse. The data engineering team loads data daily from Amazon S3 using COPY commands. Recently, the load performance has degraded because the S3 bucket contains many small files. The team needs to optimize the COPY operation to improve performance. Which approach should they take?
122A data engineering team is using Amazon EMR to process large datasets stored in Amazon S3. The cluster uses Spot Instances for cost savings. During processing, the team notices that tasks are failing due to Spot Instance interruptions. The team needs to make the EMR job resilient to Spot interruptions without increasing costs significantly. Which solution should they implement?
123A data engineering team is using AWS Glue to catalog data in an S3 data lake. They have a Glue crawler that runs daily to update the Data Catalog. Recently, they noticed that the crawler is taking longer to run and sometimes fails because of a timeout. The team suspects the issue is due to the large number of small files in the S3 bucket. They need to improve crawler performance and reliability. Which solution should they implement?
124A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is JSON and each record is about 2 KB. The delivery stream is configured to buffer incoming data to 5 MB or 60 seconds, whichever comes first. The data engineering team notices that the S3 bucket contains many small files (average 2 MB), which makes subsequent processing inefficient. They need to reduce the number of small files without increasing the latency beyond 5 minutes. Which solution should they implement?
125A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB and has a 24/7 uptime requirement. The migration must minimize downtime. Which approach should be used?
126A company is using Amazon DynamoDB with on-demand capacity for a gaming application. During a new game launch, write traffic spikes to 50,000 writes per second, but the application experiences throttling. The DynamoDB table has a partition key of 'game_id' and a sort key of 'timestamp'. What is the MOST likely cause of throttling?
127A company uses Amazon Redshift for data warehousing. The data engineering team notices that queries are slow due to high disk I/O. The team wants to improve query performance without changing the cluster configuration. Which action should the team take?
128A company wants to use Amazon Redshift Spectrum to query data in Amazon S3. The data is in Parquet format and partitioned by date. Which step is required to enable Redshift Spectrum?
129A company is using Amazon RDS for PostgreSQL with Multi-AZ deployment. The primary instance fails and a failover occurs. After the failover, the application cannot connect to the database. What is the MOST likely cause?
130A company is using Amazon ElastiCache for Redis to cache frequently accessed data. The cache hit ratio is low, and the engineering team suspects that the eviction policy is causing important data to be removed. Which eviction policy should be used to minimize eviction of the most frequently accessed keys?
131A company is designing a data lake on Amazon S3 for analytics. The data includes sensitive personally identifiable information (PII). Which TWO actions should the company take to protect the data? (Choose TWO.)
132An IAM role 'DataLakeRole' has the above S3 bucket policy attached to an S3 bucket. The role is assumed by an AWS Glue job. The Glue job is failing with 'Access Denied' errors when trying to list objects in the bucket. Which action should be added to the policy to fix the issue?
133A data engineer needs to store JSON documents that are frequently updated and require ACID transactions. Which AWS database service is most appropriate?
134A company is using Amazon RDS for MySQL and needs to reduce read latency for a global user base. Which AWS feature should be implemented?
135A company has an Amazon S3 bucket with versioning enabled. They want to automatically delete noncurrent versions of objects after 30 days. Which lifecycle rule action should be used?
136A data engineer is designing a data lake on Amazon S3. The data is frequently accessed by multiple analytics services, and the company needs to enforce fine-grained access control based on data tags. Which combination of AWS services should be used?
137Which TWO of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose TWO.)
138Which THREE of the following are valid storage classes in Amazon S3? (Choose THREE.)
139Which TWO of the following are best practices for Amazon Redshift table design? (Choose TWO.)
140A data engineer needs to store JSON documents that are frequently read and written by a web application. The data has a flexible schema and requires low-latency queries on primary key lookups. Which AWS service is MOST suitable?
141A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources in Parquet format, and the schema evolves over time. Which approach allows querying the data with Amazon Athena while supporting schema evolution?
142A company needs to store relational data that requires complex joins and transactional consistency. The workload is predictable and the data size is less than 500 GB. Which AWS service is MOST cost-effective for this use case?
143A company uses Amazon DynamoDB for a gaming leaderboard. The table has a partition key of 'GameId' and a sort key of 'Score'. The application needs to query the top 10 scores for a given game. Which DynamoDB feature should be used for optimal performance?
144A company is migrating an on-premises Hadoop cluster to AWS. The cluster processes large files in CSV format using Apache Spark. Which data store should be used as the primary storage for the data lake to optimize cost and performance?
145A company has an Amazon Redshift cluster with a mix of frequently accessed hot data and rarely accessed cold data. They want to reduce storage costs without affecting query performance for the hot data. Which strategy is MOST effective?
146A company needs to store archival logs that must be retained for 10 years. The logs are accessed infrequently, but when accessed, retrieval must occur within 12 hours. Which storage class is MOST cost-effective?
147Refer to the exhibit. A data engineer configured the lifecycle policy shown. The 'logs/' prefix contains important audit logs. After 365 days, what happens to the objects?
148Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to decrypt objects in an S3 bucket using a customer managed KMS key. What is the effect of this policy?
149Refer to the exhibit. A data engineer notices that the Redshift cluster 'mycluster' does not have automated backups beyond 7 days. However, the compliance team requires a minimum of 35 days of backup retention. What should the engineer do?
150A company stores sensitive data in an S3 bucket. To meet compliance requirements, they must ensure that all objects are encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement should be applied to deny uploads that do not use the required encryption?
151A data engineer needs to migrate an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB and has a continuous stream of write operations. The migration should minimize downtime. Which AWS service should be used?
152A company runs a critical application on Amazon RDS for MySQL. To ensure high availability and automatic failover, the database is deployed as a Multi-AZ DB instance. The application uses read-heavy workloads. Which additional configuration should be used to offload read traffic without impacting write performance?
153A data engineer is troubleshooting an Amazon Redshift cluster that has been experiencing slow query performance. The engineer checks the system tables and finds that many queries are waiting on 'wlm_queued' time. The cluster has 10 nodes and uses automatic WLM. What is the most likely cause?
154A company uses Amazon S3 to store large datasets. The data engineering team needs to provide access to specific objects in the bucket to external partners using presigned URLs. Each URL should expire after 12 hours. The team wants to ensure that the presigned URLs cannot be used to access other objects in the bucket. Which approach should be taken?
155Which TWO of the following are features of Amazon RDS Multi-AZ deployments? (Choose 2.)
156A data engineer runs the above command and gets the output. What does the 'MFADelete' setting imply?
157A company uses Amazon S3 to store sensitive data. The security team wants to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted at rest using server-side encryption with AWS KMS managed keys (SSE-KMS). Which bucket policy statement should be added to enforce this requirement?
158A company runs a MySQL database on Amazon RDS. The database size is 500 GB and is experiencing high read traffic. The team wants to improve read performance with minimal operational overhead. Which action should they take?
159A data engineer needs to transfer 10 TB of data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited to 100 Mbps, and the transfer must be completed within 48 hours. Which solution meets the requirements?
160A data engineer needs to store JSON documents that are frequently accessed by a low-latency web application. The data does not require complex queries, and the access pattern is primarily by a key. Which AWS service is most appropriate?
161A startup is building a mobile application that requires a database to store user profiles and preferences. The database must scale automatically with minimal administration. Which AWS service should they use?
162Which TWO features of Amazon S3 help protect data from accidental deletion or modification? (Choose two.)
163A company is using Amazon S3 to store large amounts of archival data. The data is accessed infrequently but must be immediately retrievable when needed. Which storage class is the most cost-effective choice?
164A data engineer needs to set up a new Amazon RDS for PostgreSQL database for a production workload. The database must be highly available and resilient to a single Availability Zone failure. Which configuration should the engineer choose?
165A company wants to migrate its on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which AWS service should be used for the migration?
166A data engineer is troubleshooting an Amazon DynamoDB table that has frequent throttling exceptions for write requests. The table has auto scaling enabled. What is the most likely cause?
167A company stores sensitive data in Amazon S3. They need to ensure that all objects are encrypted at rest. Which approach meets this requirement with minimal effort?
168A data engineer is designing a data lake on Amazon S3. Which feature should be used to manage the lifecycle of objects and move them to cheaper storage classes automatically?
169A company has an Amazon RDS for MySQL database that is experiencing performance issues due to a large number of read requests. The application is read-heavy and can tolerate eventually consistent reads. Which action will reduce the load on the primary database with the least operational overhead?
170A data engineer needs to store JSON documents that are accessed by a key-value pattern. The workload requires single-digit millisecond latency at any scale. Which AWS service is most appropriate?
171A company wants to enforce that all data in an S3 bucket is encrypted at rest using AWS KMS. Which bucket policy condition key should be used?
172A data engineer is designing a disaster recovery strategy for an Amazon RDS for PostgreSQL database. The primary database is in us-east-1. Which TWO approaches provide cross-region disaster recovery?
173A company is migrating a large Oracle data warehouse to Amazon Redshift. Which THREE considerations are important for optimizing the Redshift cluster?
174A data engineer is setting up Amazon S3 event notifications to trigger an AWS Lambda function when new objects are uploaded. Which TWO actions are required to enable this?
175The exhibit shows an S3 bucket policy. What is the effect of this policy?
176A data engineer notices that an Amazon Redshift cluster is running low on disk space. The cluster has three nodes of type dc2.large. Which action will increase the available storage capacity?
177A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data should be stored cost-effectively. Which storage solution meets these requirements?
178A company uses Amazon RDS for MySQL with Multi-AZ deployment. The primary instance fails, and automatic failover occurs. After failover, the application experiences higher latency. What is the most likely cause?
179A data engineer is migrating an on-premises Apache HBase workload to Amazon DynamoDB. The HBase table has a row key with composite structure: customer_id (10 chars) + timestamp (10 digits). The access pattern is to query by customer_id and retrieve the latest entries. How should the DynamoDB table be designed to optimize performance?
180A data engineer is using Amazon Athena to query data stored in an S3 bucket. The queries are running slowly. Which THREE actions can improve query performance?
181A company is building a data pipeline that ingests streaming data from IoT devices. The data must be stored in a durable, scalable, and cost-effective manner for batch processing. Which TWO AWS services should be used together?
182A data engineer applies the following IAM policy to an IAM user: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::example-bucket/*", "Condition": { "StringEquals": { "s3:x-amz-server-side-encryption": "AES256" } } } ] } ``` The user attempts to download an object from the bucket 'example-bucket' that is encrypted with SSE-S3 (AES256). Will the request succeed?
183An application uses the 'orders' DynamoDB table with the schema and provisioned throughput shown in the exhibit. The application frequently queries by customer_id (range key) without specifying the order_id (partition key). What is the most likely impact on performance?
184A data engineer runs the above SQL commands on an Amazon Redshift cluster. The table 'users' is created with DISTSTYLE EVEN. What is the effect of the DISTSTYLE EVEN on query performance?
185A company is using an Amazon RDS for MySQL database for an e-commerce application. During a sales event, the database experiences high read traffic, causing slow query performance. The company wants to reduce the read load on the primary database without changing the application code. Which solution meets these requirements?
186A company uses Amazon DynamoDB as the primary data store for a gaming application. The application experiences sudden spikes in traffic. The data engineer notices that write requests are throttled during peak times. The partition keys are well-distributed. What should the data engineer do to reduce throttling?
187A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The data engineer needs to ensure minimal downtime during migration. Which AWS service should be used to facilitate the migration?
188A company is designing a data lake on Amazon S3. The data includes personal identifiable information (PII). The data engineer must ensure that only authorized users can access the data, and that access is logged for auditing. Which combination of services should the data engineer use?
189A company is storing large amounts of log data in Amazon S3. The data is accessed frequently for the first 30 days, then rarely after that. The company wants to automatically transition the data to a lower-cost storage class after 30 days. Which S3 feature should the data engineer use?
190A company is using Amazon DynamoDB to store session data for a web application. The data engineer needs to ensure that the data is encrypted at rest. Which action should the data engineer take?
191A company is using Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key. The data engineer must ensure that only a specific IAM role can decrypt the data. Which policy should the data engineer attach to the KMS key?
192Which TWO statements are true about Amazon Redshift distribution styles? (Choose TWO.)
193Which THREE actions can help improve read performance in Amazon DynamoDB? (Choose THREE.)
194Which TWO are benefits of using Amazon S3 Object Lock? (Choose TWO.)
195A data engineer is reviewing an IAM policy that controls access to an S3 bucket. The policy is attached to a user group. The policy includes a condition that explicitly requires server-side encryption with SSE-S3 for all GetObject requests. The engineer notices that users are unable to download objects from the bucket. What is the likely cause?
196A company is using Amazon RDS for MySQL with Multi-AZ deployment. The database size is 2 TB and the workload is read-heavy. To improve read performance, which option should be used?
197A data engineer needs to store large amounts of data that is accessed infrequently but must be retrieved immediately when needed. Which Amazon S3 storage class is most cost-effective?
198A company needs to migrate an on-premises 10 TB PostgreSQL database to Amazon RDS for PostgreSQL with minimal downtime. Which AWS service should be used for the migration?
199A company needs to store JSON documents that are accessed by a key-value pattern. The data is 500 GB and requires single-digit millisecond latency. Which AWS database is most suitable?
200Which TWO actions can reduce the cost of an Amazon S3 bucket that stores infrequently accessed data? (Choose 2.)
201Which THREE steps are recommended for migrating an on-premises Oracle database to Amazon RDS for Oracle with minimal downtime? (Choose 3.)
202A data engineer needs to migrate an on-premises Apache Hadoop cluster to AWS. The cluster stores data in HDFS and runs MapReduce jobs. The company wants to minimize operational overhead and leverage serverless technologies where possible. Which AWS service should the data engineer use to replace HDFS storage?
203A data engineer is designing a data lake on AWS using Amazon S3. The data consists of CSV files generated by IoT devices. The data is accessed by multiple analytics jobs, and the engineer needs to ensure that new files are immediately visible to all consumers after writing. What S3 consistency model applies?
204A company has an Amazon RDS for PostgreSQL DB instance with a large table that is frequently updated. The data engineer needs to reduce storage costs by archiving old records that are no longer accessed. The archived records must be retained for 7 years due to compliance requirements. Which solution is MOST cost-effective?
205A data engineer is designing a multi-region disaster recovery plan for an Amazon DynamoDB table. The table stores critical user profile data and must have a Recovery Point Objective (RPO) of less than 1 minute and a Recovery Time Objective (RTO) of less than 5 minutes. Which solution meets these requirements?
206A company uses Amazon S3 to store customer documents. The data engineer needs to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted with a customer-managed AWS KMS key. What should the data engineer do?
207A company stores sensitive financial data in an Amazon Redshift cluster. The data engineer must ensure that all queries are logged for audit purposes and that the logs are stored in Amazon S3 with server-side encryption. Which THREE steps should the data engineer take to meet these requirements?
208A data engineer is optimizing an Amazon RDS for MySQL database that experiences high write throughput. The engineer wants to improve write performance and reduce latency. Which TWO database-level configuration changes can help achieve this?
209A company is using Amazon RDS for MySQL with Multi-AZ deployment. The primary DB instance experiences a hardware failure, causing automatic failover to the standby. After the failover, the application reports that the database endpoint is unreachable for about 60 seconds. What is the MOST likely cause?
210A company stores critical financial data in Amazon DynamoDB. To meet compliance requirements, the data must be encrypted at rest with a customer-managed key. Which solution should the data engineer implement?
211A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow and the system is experiencing high disk usage. The engineer suspects that the distribution style is suboptimal. Which action should the engineer take to improve query performance?
212A data engineer needs to store semi-structured JSON data from IoT devices. The data is written frequently and read occasionally. Which AWS service is MOST cost-effective for this use case?
213A company is using Amazon S3 to store sensitive customer data. The security team requires that all data be encrypted in transit and at rest. Additionally, they want to prevent any accidental public access. Which combination of actions should the data engineer take?
214A data engineer is designing a data store for a real-time leaderboard application that requires sub-millisecond read and write latency. The leaderboard stores scores for millions of users and needs to be sorted by score. Which AWS service should the engineer use?
215A data engineer needs to store archival data that is rarely accessed but must be retained for 7 years. The data should be retrievable within 12 hours. Which Amazon S3 storage class is MOST cost-effective?
216Which TWO actions can help improve the read performance of an Amazon DynamoDB table that is experiencing throttling? (Choose two.)
217Which THREE factors should a data engineer consider when choosing between Amazon RDS and Amazon DynamoDB for a new application? (Choose three.)
218Which TWO methods can be used to enforce least-privilege access to an Amazon S3 bucket? (Choose two.)
219Refer to the exhibit. A data engineer applies this bucket policy to an S3 bucket. A user within the 10.0.0.0/24 IP range attempts to upload an object to the bucket using an HTTP (non-HTTPS) request. What is the outcome?
220Refer to the exhibit. A data engineer runs the above CLI command and sees the output. The security team requires that the RDS instance not be accessible from the internet. Which change should the engineer make?
221Refer to the exhibit. A data engineer creates an Amazon Redshift table with the above DDL. The engineer runs a query to find all orders for a specific customer within a date range. Which statement about query performance is correct?
222A company uses an Amazon RDS for MySQL DB instance with Multi-AZ deployment. The primary DB instance fails unexpectedly. What happens to the database endpoint?
223A data engineer is designing a data lake on Amazon S3. Data is ingested from multiple sources in JSON format. The engineer needs to optimize query performance for Amazon Athena while minimizing storage costs. Which storage strategy should the engineer use?
224A company uses Amazon DynamoDB with on-demand capacity for a gaming application that experiences unpredictable traffic spikes. The application reads the same set of 'hot' items frequently. Users report high latency during peak hours. Which action would MOST effectively reduce read latency for the hot items?
225A data engineer is migrating an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB in size. The engineer needs to minimize downtime. Which AWS service should be used for the migration?
226A data engineer needs to store semi-structured JSON data that is accessed infrequently but requires immediate retrieval when needed. The data must be durable and cost-effective. Which Amazon S3 storage class should be used?
227A company uses Amazon DynamoDB with global tables in three AWS Regions. The data engineer needs to ensure that writes to the table in us-east-1 are replicated to other regions with minimal latency. Which DynamoDB feature should be used?
228Which TWO options are valid ways to reduce storage costs for an Amazon S3 data lake that stores historical data rarely accessed after 30 days? (Choose TWO.)
229Which TWO AWS services can be used to automatically back up an Amazon RDS for SQL Server DB instance? (Choose TWO.)
230A data engineer needs to store semi-structured data (JSON logs) from thousands of IoT devices. The data must be schema-less, highly scalable, and support low-latency queries by device ID and timestamp. Which AWS service should the engineer use?
231A company stores sensitive data in an Amazon S3 bucket. A compliance requirement mandates that all data must be encrypted at rest with a key that is automatically rotated every year. The company also needs to maintain an audit trail of who used the key. Which solution meets these requirements?
232A company runs an Amazon Redshift cluster with 10 RA3 nodes. The data warehouse stores 50 TB of data. The company notices that queries are slow and the cluster's storage utilization is high. The data engineer needs to improve query performance and reduce storage costs without changing the cluster's node count. Which action should the engineer take?
233A data engineer is designing a data lake on Amazon S3. The data lake will store raw data, transformed data, and curated datasets. The engineer needs to ensure that raw data is immutable (never overwritten or deleted) and that only authorized users can access the transformed data. Which combination of S3 features should the engineer use?
234A company runs an Amazon RDS for PostgreSQL instance that stores financial data. The company requires point-in-time recovery (PITR) with a retention period of 35 days. Additionally, the company needs to create a new database from a specific snapshot every night for testing. Which combination of actions should the data engineer take to meet these requirements?
235A media company stores video files in an S3 bucket. The files are processed by a fleet of EC2 instances that read the files, add watermarks, and write the output back to the same bucket. Recently, the processing jobs have been failing with '500 Internal Server Error' and '503 Slow Down' errors. The data engineer checks the S3 bucket metrics and sees that the PUT/GET request rate is consistently above 5,500 requests per second for a single prefix. The engineer needs to resolve the errors with minimal changes to the application code. Which course of action should the engineer take?
236A company stores its application logs in an Amazon S3 bucket. The logs are accessed frequently for the first 30 days, after which they are rarely accessed but must be retained for 7 years for compliance. The company wants to optimize storage costs while maintaining immediate retrieval availability for the first 30 days and the ability to retrieve logs within 12 hours after that. Which lifecycle policy should the data engineer configure?
237A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB in size. The migration must have minimal downtime and must be completed within a week. Which AWS service should the data engineer use to perform the migration?
238A data engineer is designing a data lake on Amazon S3 that will store sensitive financial data. The engineer needs to implement encryption at rest and ensure that only authorized users can access the data. Which TWO actions should the engineer take to meet these requirements? (Choose TWO.)
239A company runs an Amazon RDS for PostgreSQL instance for an OLTP application. The database size is 500 GB. The company wants to minimize downtime during backups and ensure point-in-time recovery (PITR) for the last 7 days. Which TWO features should the company use? (Choose TWO.)
240A company uses AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by date. The company wants to improve query performance in Amazon Athena and reduce costs. Which THREE actions should the company take? (Choose THREE.)
241A large e-commerce company uses Amazon DynamoDB to store shopping cart data. The table has a partition key of 'user_id' and a sort key of 'item_id'. The application performs frequent updates to the 'quantity' attribute for items in a user's cart. Recently, the operations team noticed that write requests are being throttled during peak shopping hours. The table is provisioned with 10,000 write capacity units (WCUs) and uses DynamoDB Accelerator (DAX) for read caching. The data engineer suspects that the throttling is due to hot partitions. The application uses a single AWS SDK client configured with retries. After reviewing the Amazon CloudWatch metrics, the engineer sees that the WriteThrottleEvents metric spikes for a few partition keys. The table has a high number of partitions. What should the data engineer do to resolve the throttling issue with minimal application changes?
242A data engineer is designing a data lake on Amazon S3. The team wants to optimize query performance and reduce storage costs for a large dataset of JSON logs that are queried frequently by Amazon Athena. The logs are currently stored as uncompressed JSON files, each around 1 GB, in a single prefix. The engineer needs to improve query performance and reduce costs without changing the data format. Which action should the engineer take?
243A data engineer manages a large Amazon S3 data lake with millions of small JSON files ingested daily. Amazon Athena queries against this data lake are slow and costly due to high per-query data scanned. The engineer wants to optimize the storage layout to improve query performance and reduce cost, while keeping the data queryable in place. Which solution should the engineer implement?
244A media company stores millions of small JSON files in an Amazon S3 bucket and queries them with Amazon Athena. Analysts report that queries scan far more data than expected, and costs are rising. The data engineer confirms that the files are uncompressed, use no partitioning, and are stored as newline-delimited JSON. Which change will MOST reduce the data scanned per query?
245A financial services company stores sensitive transaction data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys be automatically rotated every year. A data engineer needs to configure the S3 bucket to meet these requirements with minimal ongoing operational effort. Which solution should the engineer implement?
246A financial analytics company stores daily transaction records in Amazon S3 as Apache Parquet files, partitioned by year/month/day. The data engineering team queries these files with Amazon Athena. To reduce query runtime and cost, they want to apply fine-grained access control and column-level filtering without changing the files. Which solution should they use?
247A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a nested folder structure. The engineer wants to automatically discover the schema and update the AWS Glue Data Catalog as new files are added. Which solution should the engineer use?
248A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a folder structure. The engineer wants to use AWS Glue crawlers to automatically infer schemas and create tables in the AWS Glue Data Catalog. The crawler should run daily to detect new files and schema changes. Which configuration should the engineer use for the crawler?
249A company is using an Amazon RDS for PostgreSQL database to store application data. The data engineering team needs to run complex analytical queries that join multiple large tables. These queries are causing performance degradation on the production database. The team wants to offload the analytical workload to a separate system that can handle large-scale data processing. Which AWS service should the team use?
250A data engineer is building an Amazon Redshift data warehouse. The cluster will ingest data from Amazon S3 using the COPY command. The engineer needs to ensure that the data is loaded in a way that maximizes query performance for future complex analytical queries. The data is currently stored as uncompressed CSV files in S3. Which action should the engineer take to optimize the load and subsequent query performance?
251A data engineer manages an Amazon DynamoDB table used for a high-traffic gaming leaderboard. The table uses on-demand capacity mode and has a partition key of UserId (string) with no sort key. The leaderboard must retrieve the top 100 scores across all users. Currently, the engineer scans the entire table and sorts the results in application code, which takes several seconds and consumes large amounts of read capacity. What should the engineer do to improve the performance of retrieving the top scores?
252A media company ingests thousands of small JSON files per hour into an Amazon S3 bucket. A data engineer needs to convert these files into a compact, columnar format for efficient querying with Amazon Athena. The engineer wants to minimize storage costs and improve query performance. Which approach should the engineer take?
253A data engineer maintains an Amazon S3 data lake with millions of small JSON objects. The engineer needs to improve query performance by reducing the number of objects and compressing them into a columnar format that Amazon Athena can query efficiently. The data must remain partitioned by date. Which solution should the engineer use?
254A data engineer is designing a solution to ingest streaming data from Amazon Kinesis Data Streams into an Amazon Redshift cluster for near-real-time analytics. The engineer needs to ensure that data is loaded efficiently and that the Redshift cluster can handle the ingestion load without impacting query performance. Which approach should the engineer use?
255A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Apache Parquet format and partitioned by year, month, and day. The engineer notices that AWS Glue crawlers are taking a long time to run and are not correctly identifying new partitions. The engineer needs to improve the crawler performance and ensure new partitions are added automatically. Which action should the engineer take?
256A company stores sensitive financial data in an Amazon S3 bucket. The data engineering team must ensure that all data is encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys are rotated annually. The team also needs to audit key usage. Which solution meets these requirements?
257A data engineer is configuring an Amazon S3 bucket to store sensitive financial data. The company requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption key be automatically rotated every year. The engineer creates a KMS customer managed key and enables automatic rotation. When uploading objects using the AWS CLI, the engineer uses the --sse aws:kms parameter but does not specify a key ID. What is the result of this configuration?
258A retail company uses Amazon DynamoDB to store product catalog data. The table has a partition key of ProductID and a sort key of Category. The company needs to retrieve all products in a specific category, sorted by ProductID. Which operation should be used?
259A data engineer is designing a data warehouse on Amazon Redshift. The workload includes many ad-hoc queries that filter on a high-cardinality column, such as customer_id, and join large dimension tables. The engineer wants to improve query performance by choosing an appropriate distribution style and sort key. Which combination should the engineer use?
260A data engineer is building a data lake on Amazon S3. The engineer needs to catalog metadata for data stored in Parquet format and make it queryable by Amazon Athena. The data is partitioned by year, month, and day in the S3 path. Which AWS service should the engineer use to create and manage the table definitions and partitions?
261A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by year, month, and day. The engineer needs to ensure that AWS Glue crawlers correctly identify the partitions and that Amazon Athena queries can efficiently prune partitions. Which action should the engineer take?
262A data engineering team stores clickstream events in an Amazon S3 bucket under the prefix s3://analytics/raw/. New objects arrive continuously, and the team wants Amazon Athena queries to scan only the events for the current day without scanning the entire prefix. The events are written as JSON files partitioned by year/month/day, but queries still scan all partitions because the partition metadata is not registered. Which action should the data engineer take to enable partition pruning in Athena?
263A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The workload includes frequent queries that retrieve all sessions for a user within a specific time range. The engineer notices that some queries are slow and wants to optimize the table design. Which action should the engineer take to improve query performance?
264A logistics company stores shipment tracking events in an Amazon DynamoDB table. The table uses a partition key of shipment_id and a sort key of event_timestamp. Analysts frequently run queries that filter by shipment_id and a range of event_timestamp values. The data engineer must ensure these queries are efficient and consume minimal read capacity. What should the data engineer do?
265A company stores application logs in Amazon S3 and uses AWS Glue crawlers to populate the AWS Glue Data Catalog. A data engineer needs to query the logs with Amazon Athena. The logs are partitioned by year/month/day in S3, but Athena queries are scanning all partitions and returning errors about missing partitions. What should the engineer do to enable partition pruning?
266A data engineer manages an Amazon S3 data lake with millions of small JSON files. To improve query performance with Amazon Athena, the engineer wants to compact these files into larger Parquet files. The engineer must also minimize ongoing storage costs. Which solution should the engineer implement?
267A data engineer is building an Amazon DynamoDB table that will store IoT sensor readings. Each reading has a deviceId (partition key) and a timestamp (sort key). The team wants to retrieve all readings for a device within the last 24 hours, and they also want to minimize the number of read capacity units consumed. Which access pattern should the engineer implement?
268A data engineer is building a data lake on Amazon S3. The engineer needs to store structured data that will be queried by Amazon Athena. The data is currently in CSV format and is partitioned by date. The engineer wants to improve query performance and reduce the amount of data scanned. Which action should the engineer take?
269A data engineer is configuring an Amazon Redshift cluster for a workload that runs large nightly ELT jobs loading data from Amazon S3 and then executes complex analytical queries. The team wants to improve query performance and reduce the time the cluster spends on data loading. Which TWO configuration choices should the engineer make? (Choose two.)
270A data engineer is designing an Amazon S3 data lake and needs to enforce schema-on-read for a dataset that is queried by Amazon Athena. The data is stored as Parquet files partitioned by year, month, and day. The engineer wants to minimize the amount of data scanned by queries that filter on a specific date range. Which approach should the engineer take?
271A data engineer manages an Amazon Redshift cluster that stores sales data in a table with a sort key on the sale_date column. The table is growing rapidly, and queries that filter by sale_date are becoming slower. The engineer notices that the table has a high percentage of unsorted rows. What should the engineer do to improve query performance with the least effort?
272A media company stores millions of thumbnail images in an Amazon S3 bucket. Analysts run ad hoc queries against the image metadata, which is kept as JSON objects in the same bucket. Query latency is unpredictable and costs are rising because Athena scans large volumes of JSON for every query. The team wants faster queries and lower scan cost while keeping the data in S3 and queryable with SQL. Which change should the data engineer make?
273A company stores application logs in Amazon S3 and needs to query them using standard SQL. The logs are in JSON format and are updated daily. The data engineering team wants a serverless solution that requires minimal management and can automatically discover the schema. Which AWS service should they use?
274A data engineer is building a data lake on Amazon S3 and must choose the optimal file format for a dataset that is queried by Amazon Athena. The queries typically select a few columns from wide tables containing hundreds of columns, and the data volume is in terabytes. The engineer wants to minimize query scan costs and improve performance. Which file format should the engineer use?
275A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query loads. The cluster uses a dc2.large node type with 2 nodes. Analysis shows that the workload involves frequent large table scans and complex joins. The engineer wants to improve query performance without changing the overall data volume. Which action should the engineer take?
276A company uses Amazon RDS for PostgreSQL to store customer data. The data engineer needs to ensure that the database can be restored to any point in time within the last 35 days. The engineer also wants to minimize the impact on the production database during backups. What should the engineer do?
277A data engineer manages an Amazon Redshift provisioned cluster that serves a nightly ELT workload. The cluster's largest fact table is loaded with new rows each night, and queries frequently filter on a date column and join to a customer dimension. The engineer wants to improve query performance and reduce the time spent vacuuming. Which TWO actions should the engineer take? (Choose two.)
278A data engineer is using Amazon Redshift and needs to improve the performance of complex queries that join large tables. The engineer has already set the distribution style to KEY on the join columns. What additional step should the engineer take to optimize the join performance?
279A data engineer is designing a data lake on Amazon S3 and needs to ensure that data is encrypted at rest. The company requires that encryption keys be managed by AWS and automatically rotated annually. The engineer also wants to audit key usage. Which S3 encryption option should the engineer choose?
280A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it using Amazon Athena. The data is generated continuously and appended to S3 in small files. The engineer wants to optimize query performance and reduce costs. Which action should the engineer take?
281A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The engineer needs to retrieve all sessions for a specific user that started within the last 30 days. Which operation should the engineer use to achieve the lowest latency?
282A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a Redshift table. The data is in Parquet format and is partitioned by date in S3. The engineer wants to load only the data for the last 7 days to reduce load time and cost. Which Redshift command should the engineer use?
283A data engineer is designing a data store for a real-time analytics application that requires low-latency reads and writes at scale. The data model includes time-series data with high ingest rates and queries that aggregate data over sliding time windows. The engineer needs a fully managed AWS service that supports automatic scaling and can handle millions of writes per second. Which service should the engineer choose?
284A data engineer manages an Amazon S3 data lake with millions of small JSON files ingested continuously. Amazon Athena queries over this data are slow and expensive because each query scans many small objects. The engineer wants to improve query performance and reduce cost without changing the data content. Which solution should the engineer implement?
285A data engineer is designing a data lake on Amazon S3 for a retail company. The company ingests point-of-sale data as small JSON files every few minutes, totaling about 5 GB per day. Analysts query the data with Amazon Athena, and costs are rising due to many small files and full scans. The engineer wants to reduce Athena query costs and improve performance while keeping the data in S3. Which TWO actions should the engineer take? (Choose two.)
286A company uses Amazon DynamoDB to store user session data. The table has a partition key of UserID and a sort key of SessionStartTime. The application frequently queries for all sessions of a specific user within a date range. The table is provisioned with 1000 RCUs and 1000 WCUs. During peak hours, the application experiences throttling on read requests. Which action should a data engineer take to resolve the throttling with minimal changes?
287A data engineer is deploying an Amazon Redshift cluster that must be accessible only from within a private VPC and must not have a public IP address. The cluster will be queried by an Amazon EMR cluster in the same VPC and by on-premises BI tools over a VPN connection. Which configuration should the engineer choose?
288A data engineer is designing an Amazon DynamoDB table for an order-processing application. The table uses a partition key of order_id and a sort key of order_date. The application needs to retrieve all orders for a specific customer within a date range, and the queries must be efficient at scale. The engineer must choose a design that supports these access patterns without full table scans. What should the engineer do?
289A data engineer is optimizing an Amazon Redshift cluster for a workload that includes frequent complex queries with multiple joins and aggregations. The engineer wants to improve query performance by using appropriate distribution styles and sort keys. Which TWO actions should the engineer take? (Choose two.)
290A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. Each item has a partition key of deviceId and a sort key of timestamp. The table is configured with on-demand capacity mode. The engineer needs to retrieve all readings for a specific device within the last 24 hours, and the query must return results sorted by timestamp in ascending order. Which operation should the engineer use?
291A data engineer is designing a data lake on Amazon S3 and needs to catalog data using the AWS Glue Data Catalog. The data is stored in Parquet format, partitioned by year/month/day. The engineer wants to query the data using Amazon Athena and ensure that partition pruning occurs to minimize query costs. Which action should the engineer take?
292A data engineer manages an Amazon S3 data lake that holds sensitive customer transaction logs. Compliance requires that all objects be encrypted at rest with keys that the company rotates every 90 days and fully controls, including the ability to immediately revoke access and audit key usage separately from other AWS accounts. The engineer must choose an encryption method that meets these requirements with minimal operational overhead. Which solution should the engineer implement?
293A data engineer is designing a DynamoDB table for an application that requires strongly consistent reads and supports a global secondary index (GSI). The engineer needs to ensure that queries on the GSI return the most up-to-date data. Which statement about DynamoDB read consistency is correct?
294A data engineer is configuring an AWS Glue ETL job to read from an Amazon S3 bucket that contains Apache Parquet files partitioned by year, month, and day. The engineer wants the job to only process data for the year 2023 and month 10, and to minimize the amount of data scanned. The Glue job uses the Glue Data Catalog table `sales_data` with the correct partition structure. What is the MOST efficient way to configure the job to read only the required partitions?
295A company stores sensitive data in Amazon S3 and needs to ensure that data is encrypted at rest. The security team requires that the company manage its own encryption keys and have the ability to audit key usage. Which S3 encryption option should the data engineer choose?
296A data engineer is responsible for an Amazon Redshift cluster that ingests data continuously from Amazon Kinesis Data Streams. The engineer needs to ensure that the raw streaming data is immediately queryable in Redshift with minimal latency. Which approach should the engineer take?
297A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. The table uses a partition key of deviceId and a sort key of timestamp, with a provisioned read capacity of 100 RCUs. During a sudden spike in traffic, the engineer observes throttling on read operations even though the consumed read capacity is well below the provisioned limit. What is the MOST likely cause of the throttling?
298A data engineer stores Apache Parquet files in an Amazon S3 data lake partitioned by dt=YYYY-MM-DD. Analysts query the data with Amazon Athena, and monthly reports that scan one month of data are slow and expensive. The engineer confirms that queries filter on the dt column. Which action will MOST effectively reduce the amount of data scanned by these reports?
299A data engineer is configuring an Amazon S3 bucket for a data lake. The bucket must store sensitive financial data and comply with a regulation that requires all data to be encrypted at rest with keys that are automatically rotated every year. The engineer also needs to audit key usage and control access to the keys separately from other AWS services. Which encryption option should the engineer choose?
300A data engineer is building a data lake on Amazon S3. The raw data arrives as JSON files, but the analytics team needs to query the data using standard SQL in Amazon Athena with optimal performance and minimal cost. The engineer wants to convert the JSON to a columnar format that supports predicate pushdown and is natively supported by Athena. Which storage format should the engineer choose?
Deep-dive questions
The most-searched questions in this domain — detailed explanations, worked examples, full answer breakdowns.
Candidates must configure S3 lifecycle policies, DynamoDB capacity and GSIs, Glue Data Catalog partitions, and Lake Formation grants. Get S3 storage class transitions and Lake Formation column-level permissions right, since most scenario questions hinge on least-privilege access and cost-optimal storage.
The Courseiva DEA-C01 question bank contains 300 questions in the Data Store Management domain, covering the 26% of the exam attributed to this domain in the official Amazon Web Services blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data Store Management domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included