DEA-C01 · domain
Data Store Management
Data Store Management is 26% of DEA-C01 and covers choosing, configuring and securing AWS storage for analytics workloads. Expect scenario questions on S3 storage classes and lifecycle rules, DynamoDB capacity and indexes, Lake Formation permissions, Glue Data Catalog tables and partitions, and KMS encryption choices across Redshift, RDS and S3.
Focused practice
Practice Data Store Management questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Data Store Management
Candidates must configure S3 lifecycle policies, DynamoDB capacity and GSIs, Glue Data Catalog partitions, and Lake Formation grants. Get S3 storage class transitions and Lake Formation column-level permissions right, since most scenario questions hinge on least-privilege access and cost-optimal storage.
Selecting S3 storage classes, lifecycle transitions and Intelligent-Tiering for cost and access patterns
Configuring DynamoDB partition keys, GSIs/LSIs, on-demand versus provisioned capacity and DynamoDB Streams
Managing Glue Data Catalog databases, tables, partitions and crawler classification of S3 data
Applying Lake Formation grants, S3 bucket policies and KMS keys for fine-grained data access control
Watch out for
Common Data Store Management exam traps
- ▸Choosing S3 Glacier Deep Archive for data needing millisecond retrieval, ignoring the hours-long restore time and retrieval charges
- ▸Using a low-cardinality DynamoDB partition key, causing hot partitions and throttling instead of even read/write distribution
- ▸Granting IAM permissions on S3 but forgetting Lake Formation table grants, so Athena and Redshift Spectrum queries still fail
Question index
All Data Store Management questions (358)
Click any question to see the full explanation, or start a practice session above.
A company stores application logs in Amazon S3 and wants to analyze them using Amazon Athena. The logs are in JSON format and are compressed with gzip. The data engineer needs to create an Athena table that can query these logs efficiently. The logs are stored in an S3 bucket with the prefix logs/year=2023/month=10/day=15/. The engineer wants to minimize query costs and improve performance. Which action should the engineer take?
Easy2A data engineer is optimizing an Amazon S3 data lake for cost and performance. The data lake contains large volumes of CSV files that are queried by Amazon Athena. The engineer wants to reduce query costs and improve query performance. Which TWO actions should the engineer take? (Choose two.)
Medium3A company stores sensitive financial data in an Amazon S3 bucket. The data engineering team must ensure that all data is encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys are rotated annually. The team also needs to audit key usage. Which solution meets these requirements?
Hard4A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is JSON and each record is about 2 KB. The delivery stream is configured to buffer incoming data to 5 MB or 60 seconds, whichever comes first. The data engineering team notices that the S3 bucket contains many small files (average 2 MB), which makes subsequent processing inefficient. They need to reduce the number of small files without increasing the latency beyond 5 minutes. Which solution should they implement?
Medium5A data engineer is building an Amazon Redshift data warehouse. The cluster will ingest data from Amazon S3 using the COPY command. The engineer needs to ensure that the data is loaded in a way that maximizes query performance for future complex analytical queries. The data is currently stored as uncompressed CSV files in S3. Which action should the engineer take to optimize the load and subsequent query performance?
Medium6Which TWO features of Amazon S3 help protect data from accidental deletion or modification? (Choose two.)
Easy7A data engineer manages an Amazon S3 data lake that ingests millions of small JSON files daily from an IoT fleet. Query performance in Amazon Athena has degraded significantly, and each query scans far more data than expected. The engineer wants to reduce per-query cost and improve performance without changing the raw data. Which solution should the engineer implement?
Medium8A data engineering team stores clickstream events in an Amazon S3 bucket under the prefix s3://analytics/raw/. New objects arrive continuously, and the team wants Amazon Athena queries to scan only the events for the current day without scanning the entire prefix. The events are written as JSON files partitioned by year/month/day, but queries still scan all partitions because the partition metadata is not registered. Which action should the data engineer take to enable partition pruning in Athena?
Medium9Which TWO of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose TWO.)
Medium10A financial services company stores trade records in Amazon DynamoDB. An application performs many reads per second for individual trades by tradeId and also runs a nightly analytics job that must read every trade for a given trading day. The table uses tradeId as the partition key only. The nightly job currently performs a full table scan and is slow and expensive. The company wants to optimize the nightly access pattern with minimal application changes. Which change should the data engineer make?
Hard11A data engineer notices that an Amazon Redshift cluster is experiencing slow query performance. The engineer suspects that tables are not properly sorted. Which diagnostic query should the engineer run to identify unsorted rows?
Medium12A company uses Amazon DynamoDB for a gaming application. They need to store player session data that expires after 24 hours. Which DynamoDB feature should they use to automatically delete expired items?
Easy13A company is using Amazon Redshift for data warehousing. The data engineer notices that the STL_ALERT_EVENT_LOG table shows many 'missing statistics' alerts. What is the best course of action to address this issue?
Hard14A company is using Amazon RDS for MySQL and needs to reduce read latency for a global user base. Which AWS feature should be implemented?
Medium15A company runs a multi-AZ Amazon RDS for PostgreSQL instance. They need to run a one-time analytical query that will take several hours and consume significant I/O. The query should not impact the primary workload. What should the data engineer do?
Medium16A data engineer is designing a multi-region disaster recovery plan for an Amazon DynamoDB table. The table stores critical user profile data and must have a Recovery Point Objective (RPO) of less than 1 minute and a Recovery Time Objective (RTO) of less than 5 minutes. Which solution meets these requirements?
Hard17A company has an Amazon RDS for PostgreSQL DB instance with a large table that is frequently updated. The data engineer needs to reduce storage costs by archiving old records that are no longer accessed. The archived records must be retained for 7 years due to compliance requirements. Which solution is MOST cost-effective?
Medium18A data engineering team is building a data lake on Amazon S3. They need to catalog data and make it queryable by Amazon Athena and Amazon Redshift Spectrum. The data arrives in multiple formats and the schema evolves frequently. Which TWO actions should the team take to support schema evolution and efficient querying? (Choose two.)
Hard19A company uses Amazon DynamoDB as the primary data store for a gaming application. The application experiences sudden spikes in traffic. The data engineer notices that write requests are throttled during peak times. The partition keys are well-distributed. What should the data engineer do to reduce throttling?
Hard20A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The migration must have minimal downtime. Which service should be used to perform the migration?
Medium21A company is storing large amounts of log data in Amazon S3. The data is accessed frequently for the first 30 days, then rarely after that. The company wants to automatically transition the data to a lower-cost storage class after 30 days. Which S3 feature should the data engineer use?
Easy22A data engineer needs to store large volumes of infrequently accessed compliance data in Amazon S3 for 10 years. The data must be retrievable within 12 hours if required for audits. The engineer wants the most cost-effective storage solution. Which S3 storage class should be used?
Easy23A data engineer is responsible for an Amazon Redshift cluster that ingests data continuously from Amazon Kinesis Data Streams. The engineer needs to ensure that the raw streaming data is immediately queryable in Redshift with minimal latency. Which approach should the engineer take?
Medium24A data engineer is configuring an Amazon S3 lifecycle policy to transition objects to S3 Glacier Deep Archive after 90 days. The bucket receives new objects daily. The engineer wants to ensure that objects are not deleted before 90 days. Which lifecycle action should be used?
Easy25A company needs to store application log files for 90 days for compliance. The logs are generated continuously and are rarely accessed after 30 days. The data engineer must minimize storage costs. Which storage solution should the engineer choose?
Easy26Which TWO actions are recommended for securing data at rest in Amazon S3? (Choose two.)
Medium27A company runs a data warehouse on Amazon Redshift. Queries are slow, and the team suspects data distribution is skewed. Which approach would best help identify distribution skew?
Medium28A data engineer is troubleshooting an Amazon DynamoDB table that has frequent throttling exceptions for write requests. The table has auto scaling enabled. What is the most likely cause?
Hard29A company is running a data warehouse on Amazon Redshift. The data engineering team notices that query performance has degraded over time. They suspect that data distribution is causing excessive data movement between nodes. The table is joined frequently on the customer_id column. Which column should be chosen as the distribution key to optimize join performance?
Medium30A data engineer created the IAM policy shown in the exhibit. The engineer then attempts to upload an object to 'my-bucket' using the AWS CLI with the command: aws s3 cp file.txt s3://my-bucket/ --sse aws:kms. The upload fails with an 'AccessDenied' error. What is the most likely cause?
Hard31A data engineer is designing a data lake on AWS using Amazon S3. The data consists of CSV files generated by IoT devices. The data is accessed by multiple analytics jobs, and the engineer needs to ensure that new files are immediately visible to all consumers after writing. What S3 consistency model applies?
Easy32A data engineer reviewed the S3 lifecycle policy shown in the exhibit. The engineer notices that objects under the 'logs/' prefix are being deleted after 365 days. The business requirement is to retain logs for at least 5 years. What should the engineer change in the lifecycle policy?
Medium33A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a nested folder structure. The engineer wants to automatically discover the schema and update the AWS Glue Data Catalog as new files are added. Which solution should the engineer use?
Medium34A company is building a data pipeline that ingests streaming data from IoT devices. The data must be stored in a durable, scalable, and cost-effective manner for batch processing. Which TWO AWS services should be used together?
Easy35A company is using Amazon S3 to store large amounts of archival data. The data is accessed infrequently but must be immediately retrievable when needed. Which storage class is the most cost-effective choice?
Medium36A company is designing a data lake on Amazon S3. Which TWO strategies improve query performance for Amazon Athena?
Medium37A company is using Amazon Redshift for its data warehouse. The data engineering team needs to improve query performance for a large fact table that is frequently joined with multiple dimension tables. Which THREE strategies should be considered?
Hard38A company stores sensitive data in Amazon S3. The security team requires encryption at rest and that the encryption keys are managed by the company using AWS KMS. The data is frequently accessed by multiple AWS services. Which THREE steps should be taken to meet these requirements?
Hard39A company uses Amazon RDS for PostgreSQL. The data engineer needs to ensure that the database is automatically backed up and that backups are retained for 35 days. What is the simplest way to achieve this?
Easy40A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow due to many small inserts. Which technique would improve write performance?
Medium41A company uses Amazon DynamoDB as its primary data store for a web application. The application experiences high latency during peak hours. The data engineer notices that the table has a large number of items with the same partition key. Which DynamoDB feature should the engineer use to improve performance?
Easy42A company needs to store JSON documents that are accessed by a key-value pattern. The data is 500 GB and requires single-digit millisecond latency. Which AWS database is most suitable?
Easy43Which TWO methods can be used to enforce least-privilege access to an Amazon S3 bucket? (Choose two.)
Easy44Which TWO actions can help improve query performance in Amazon Redshift? (Choose two.)
Medium45A data engineering team is designing a data lake on Amazon S3. They need to store raw data in its original format and transformed data in Parquet. The data is accessed by multiple analytics services, including Amazon Athena and Amazon Redshift Spectrum. Compliance requirements mandate that all data be encrypted at rest with AWS KMS and that the encryption keys be rotated every 90 days. Which S3 bucket configuration meets these requirements?
Hard46A data engineer is designing a data lake on Amazon S3. The team wants to optimize query performance and reduce storage costs for a large dataset of JSON logs that are queried frequently by Amazon Athena. The logs are currently stored as uncompressed JSON files, each around 1 GB, in a single prefix. The engineer needs to improve query performance and reduce costs without changing the data format. Which action should the engineer take?
Medium47A company runs a real-time analytics platform on Amazon ECS that ingests streaming data from Amazon Kinesis Data Streams, processes it, and stores results in Amazon DynamoDB. The data volume spikes unpredictably, causing DynamoDB to throttle write requests. The application uses on-demand capacity mode. The data engineer notices that the throttling occurs on a specific partition due to a hot key. The hot key is a customer ID that receives a disproportionate number of writes. The application cannot change the partition key design immediately. The engineer needs to reduce throttling while maintaining low latency. Which solution is most effective?
Hard48A company uses Amazon DynamoDB as the primary data store for a web application. The application experiences occasional throttling on write requests. The data engineer needs to implement a solution that handles throttling gracefully without losing data. Which approach should the engineer use?
Easy49A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for later analysis using Amazon Athena. The logs are generated continuously, and the total volume is about 1 TB per day. The data must be queryable within minutes of arrival. Which storage solution is most appropriate?
Easy50A data engineer manages an Amazon S3 data lake where analytics queries run through Amazon Athena. Monthly partition folders hold Parquet files, and each partition contains tens of thousands of small files averaging 40 KB. Athena queries that scan a single month take much longer than expected and consume far more bytes scanned than the actual data volume. The engineer must improve query performance without changing the table schema or the folder layout. What should the engineer do?
Medium51A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data is immutable and must be stored cost-effectively. Which AWS service should the engineer use?
Easy52A data engineer manages an Amazon Redshift cluster that stores sales data in a table with a sort key on the sale_date column. The table is growing rapidly, and queries that filter by sale_date are becoming slower. The engineer notices that the table has a high percentage of unsorted rows. What should the engineer do to improve query performance with the least effort?
Medium53Which THREE of the following are valid storage classes in Amazon S3? (Choose THREE.)
Medium54A data engineer is designing a disaster recovery strategy for an Amazon RDS for PostgreSQL database. The primary database is in us-east-1. Which TWO approaches provide cross-region disaster recovery?
Medium55A company is using Amazon ElastiCache for Redis to cache frequently accessed data. The cache hit ratio is low, and the engineering team suspects that the eviction policy is causing important data to be removed. Which eviction policy should be used to minimize eviction of the most frequently accessed keys?
Hard56A data engineer is designing an Amazon DynamoDB table for an order-processing application. The table uses a partition key of order_id and a sort key of order_date. The application needs to retrieve all orders for a specific customer within a date range, and the queries must be efficient at scale. The engineer must choose a design that supports these access patterns without full table scans. What should the engineer do?
Hard57A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query loads. The cluster uses a dc2.large node type with 2 nodes. Analysis shows that the workload involves frequent large table scans and complex joins. The engineer wants to improve query performance without changing the overall data volume. Which action should the engineer take?
Hard58Which THREE actions can help improve read performance in Amazon DynamoDB? (Choose THREE.)
Easy59A data engineer is building a data lake on Amazon S3. The raw data arrives as JSON files, but the analytics team needs to query the data using standard SQL in Amazon Athena with optimal performance and minimal cost. The engineer wants to convert the JSON to a columnar format that supports predicate pushdown and is natively supported by Athena. Which storage format should the engineer choose?
Medium60A data engineer manages an Amazon S3 data lake with millions of small JSON files under prefixes partitioned by year/month/day. Amazon Athena queries scan far more data than expected and return slowly. The engineer wants to reduce bytes scanned and improve query performance while keeping files in S3 and queryable by Athena. Which solution meets these requirements with the LEAST operational overhead?
Medium61A data engineer maintains an Amazon S3 data lake with millions of small JSON objects. The engineer needs to improve query performance by reducing the number of objects and compressing them into a columnar format that Amazon Athena can query efficiently. The data must remain partitioned by date. Which solution should the engineer use?
Medium62A data engineer needs to store semi-structured JSON event logs in a data lake on Amazon S3 and query them with Amazon Athena using SQL, including filtering on individual JSON attributes. The team wants to avoid transforming the files before querying. Which approach should the engineer use?
Easy63A company is designing a data lake on Amazon S3. The data includes CSV files, Parquet files, and images. The data engineering team needs to catalog the metadata and enable SQL queries. Which TWO AWS services should be used together?
Easy64An e-commerce company uses Amazon DynamoDB as the primary data store for its product catalog. The table has a simple primary key (ProductID) and handles 10,000 writes per second during peak hours. Recently, the engineering team noticed increased write latency and throttled requests during peak times. The table's provisioned write capacity is set to 12,000 WCU. What is the most likely cause of the throttling?
Medium65A media company ingests millions of small JSON files per day into an Amazon S3 bucket. Analysts run Amazon Athena queries over this data and report that each query scans far more data than the files matching their filters, resulting in high cost and slow performance. The files are partitioned by year/month/day in S3. What should a data engineer do to reduce the data scanned per query?
Hard66A company uses Amazon Redshift for its data warehouse. The data engineer notices that queries are slow on a large table that is frequently filtered on a column 'transaction_date'. Which optimization technique best improves query performance?
Hard67A company stores log files in Amazon S3. They want to automatically move logs older than 90 days to S3 Glacier Deep Archive to reduce costs. Which S3 feature should be used?
Medium68A data engineer needs to store semi-structured JSON data that is accessed infrequently but requires immediate retrieval when needed. The data must be durable and cost-effective. Which Amazon S3 storage class should be used?
Easy69A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. The table uses a partition key of deviceId and a sort key of timestamp, with a provisioned read capacity of 100 RCUs. During a sudden spike in traffic, the engineer observes throttling on read operations even though the consumed read capacity is well below the provisioned limit. What is the MOST likely cause of the throttling?
Medium70Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to decrypt objects in an S3 bucket using a customer managed KMS key. What is the effect of this policy?
Hard71A data engineer needs to store event data from IoT devices that arrives in bursts. The data is key-value and requires single-digit millisecond read and write latency. The engineer also needs to run complex analytical queries on the data for reporting. Which TWO services should be used together? (Choose TWO.)
Medium72A media company ingests thousands of small JSON files per hour into an Amazon S3 bucket. A data engineer needs to convert these files into a compact, columnar format for efficient querying with Amazon Athena. The engineer wants to minimize storage costs and improve query performance. Which approach should the engineer take?
Hard73Which TWO options are valid ways to reduce storage costs for an Amazon S3 data lake that stores historical data rarely accessed after 30 days? (Choose TWO.)
Medium74A data engineer is designing a data store for a real-time analytics application that requires low-latency reads and writes at scale. The data model includes time-series data with high ingest rates and queries that aggregate data over sliding time windows. The engineer needs a fully managed AWS service that supports automatic scaling and can handle millions of writes per second. Which service should the engineer choose?
Medium75A company's analytics team needs a petabyte-scale, fully managed data warehouse that supports standard SQL, columnar storage, and massively parallel query execution, and it must integrate with existing business intelligence tools with minimal operational effort. Which AWS service should the data engineer choose?
Easy76A data engineer needs to migrate an on-premises Apache Hadoop cluster to AWS. The cluster stores data in HDFS and runs MapReduce jobs. The company wants to minimize operational overhead and leverage serverless technologies where possible. Which AWS service should the data engineer use to replace HDFS storage?
Medium77A data engineer stores application logs in an Amazon S3 bucket. Compliance requires that log objects be retained for seven years and that they cannot be deleted or overwritten by any user, including the account root user, during that period. The engineer must configure the bucket to enforce this. Which combination of settings should the engineer apply?
Medium78A company uses Amazon S3 to store sensitive financial data. The security team requires that all objects be encrypted at rest using AWS KMS with a customer-managed key. Additionally, they want to audit all KMS decrypt calls for compliance. Which configuration should be used to meet these requirements?
Hard79A startup is building a mobile application that requires a database to store user profiles and preferences. The database must scale automatically with minimal administration. Which AWS service should they use?
Easy80Which TWO of the following are best practices for Amazon Redshift table design? (Choose TWO.)
Hard81A company is migrating a large Oracle data warehouse to Amazon Redshift. Which THREE considerations are important for optimizing the Redshift cluster?
Hard82A data engineer is building a data lake on Amazon S3 and needs to store JSON logs that will be queried by Amazon Athena. The engineer wants to minimize query cost and improve performance by reducing the amount of data scanned. The logs are approximately 1 KB each and arrive continuously. Which solution should the engineer implement?
Medium83A financial analytics company stores daily transaction records in Amazon S3 as Apache Parquet files, partitioned by year/month/day. The data engineering team queries these files with Amazon Athena. To reduce query runtime and cost, they want to apply fine-grained access control and column-level filtering without changing the files. Which solution should they use?
Medium84A data engineer is building a near-real-time ingestion pipeline into Amazon S3. Small JSON files arrive continuously from thousands of devices, and the engineer must optimize the data lake for downstream Amazon Athena queries while minimizing storage cost and query latency. Which TWO actions should the engineer take? (Choose two.)
Hard85An IAM role 'DataLakeRole' has the above S3 bucket policy attached to an S3 bucket. The role is assumed by an AWS Glue job. The Glue job is failing with 'Access Denied' errors when trying to list objects in the bucket. Which action should be added to the policy to fix the issue?
Hard86A company stores sensitive financial data in an Amazon Redshift cluster. The data engineer must ensure that all queries are logged for audit purposes and that the logs are stored in Amazon S3 with server-side encryption. Which THREE steps should the data engineer take to meet these requirements?
Hard87A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow and the system is experiencing high disk usage. The engineer suspects that the distribution style is suboptimal. Which action should the engineer take to improve query performance?
Hard88A data engineer is designing a multi-Region disaster recovery solution for an Amazon DynamoDB table. The table must be available in a secondary Region with minimal data loss and automatic failover. Which feature should be used?
Hard89A company is using an Amazon RDS for MySQL database for its e-commerce platform. During a recent flash sale, the database experienced high read traffic, causing slow query performance. The company needs a solution that offloads read traffic with minimal application changes. Which action should be taken?
Medium90Which TWO actions can help optimize Amazon S3 storage costs for a data lake? (Choose two.)
Medium91A company uses Amazon S3 to store customer documents. The data engineer needs to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted with a customer-managed AWS KMS key. What should the data engineer do?
Easy92A data engineer is troubleshooting an Amazon Redshift cluster that is running out of disk space. The engineer runs STV_PARTITIONS and notices that some slices have significantly more data than others. What is the most likely cause and solution?
Hard93A company uses Amazon DynamoDB to store user session data. The table has a partition key of UserID and a sort key of SessionStartTime. The application frequently queries for all sessions of a specific user within a date range. The table is provisioned with 1000 RCUs and 1000 WCUs. During peak hours, the application experiences throttling on read requests. Which action should a data engineer take to resolve the throttling with minimal changes?
Easy94A data engineer manages a large Amazon S3 data lake with millions of small JSON files ingested daily. Amazon Athena queries against this data lake are slow and costly due to high per-query data scanned. The engineer wants to optimize the storage layout to improve query performance and reduce cost, while keeping the data queryable in place. Which solution should the engineer implement?
Medium95A company stores sensitive data in an S3 bucket. To meet compliance requirements, they must ensure that all objects are encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement should be applied to deny uploads that do not use the required encryption?
Medium96A company is using Amazon S3 to store sensitive data. They need to ensure that all objects are encrypted at rest. Which combination of actions should be taken? (Choose TWO.)
Medium97A data engineer needs to store semi-structured JSON logs from multiple sources in a centralized data store for querying using SQL. The logs are immutable and need to be retained for 90 days. Which AWS service should be used?
Easy98A company has an Amazon S3 bucket with versioning enabled. They want to automatically delete noncurrent versions of objects after 30 days. Which lifecycle rule action should be used?
Medium99A company wants to store data from thousands of IoT devices with varying data rates. The data must be stored in a schema-on-read fashion and support SQL queries. Which AWS service should be used?
Easy100Which TWO of the following are features of Amazon RDS Multi-AZ deployments? (Choose 2.)
Easy101A data engineer is designing a data store for a real-time leaderboard application that requires sub-millisecond read and write latency. The leaderboard stores scores for millions of users and needs to be sorted by score. Which AWS service should the engineer use?
Medium102Which THREE steps are recommended for migrating an on-premises Oracle database to Amazon RDS for Oracle with minimal downtime? (Choose 3.)
Hard103A data engineer applies the bucket policy shown in the exhibit to an S3 bucket. The bucket contains sensitive data that must be encrypted at rest and accessed only over HTTPS. Which of the following statements is true?
Medium104A data engineer needs to keep a near-real-time copy of an Amazon DynamoDB table in Amazon S3 for analytics, capturing every item-level change with the before and after images and no impact on table write latency. Which approach meets these requirements with the LEAST operational effort?
Medium105A company stores its application logs in an Amazon S3 bucket. The logs are accessed frequently for the first 30 days, after which they are rarely accessed but must be retained for 7 years for compliance. The company wants to optimize storage costs while maintaining immediate retrieval availability for the first 30 days and the ability to retrieve logs within 12 hours after that. Which lifecycle policy should the data engineer configure?
Easy106A data engineer is designing a real-time analytics solution using Amazon DynamoDB. The workload requires capturing all changes to a DynamoDB table and processing them in near-real-time to update a materialized view in Amazon Redshift. Which approach should the engineer use to capture and process the changes?
Hard107A financial services company stores transaction records in an Amazon DynamoDB table. An audit requires that all data older than 7 years be automatically and permanently deleted. The data engineering team must implement this with minimal operational overhead and no application code changes. What should the team do?
Medium108A data engineer is building an Amazon Redshift data warehouse that ingests large staged files from Amazon S3 using the COPY command. The team wants to maximize load performance and minimize the time spent on ingestion. Which TWO practices should the engineer apply? (Choose two.)
Medium109A data engineer is migrating an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB in size. The engineer needs to minimize downtime. Which AWS service should be used for the migration?
Easy110A data engineer is optimizing an Amazon Redshift cluster for a workload that includes frequent complex queries with multiple joins and aggregations. The engineer wants to improve query performance by using appropriate distribution styles and sort keys. Which TWO actions should the engineer take? (Choose two.)
Hard111A company has an Amazon Redshift cluster with a mix of frequently accessed hot data and rarely accessed cold data. They want to reduce storage costs without affecting query performance for the hot data. Which strategy is MOST effective?
Hard112A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The workload includes frequent queries that retrieve all sessions for a user within a specific time range. The engineer notices that some queries are slow and wants to optimize the table design. Which action should the engineer take to improve query performance?
Hard113A company uses Amazon Redshift for a data warehouse. They notice that queries are slow due to heavy data skew. Which optimization technique should be applied first?
Hard114A data engineer is designing a data lake on Amazon S3 to store JSON logs from an application. The logs are written once and never modified. The engineer needs to query the data using Amazon Athena with the best performance and lowest cost. The engineer wants to partition the data by year, month, and day based on the log timestamp. Which approach should the engineer use to organize the S3 objects?
Medium115A data engineer needs to transfer 10 TB of data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited to 100 Mbps, and the transfer must be completed within 48 hours. Which solution meets the requirements?
Medium116Which THREE of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose three.)
Hard117A company wants to migrate its on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which AWS service should be used for the migration?
Easy118A data engineer is migrating an on-premises Apache HBase workload to Amazon DynamoDB. The HBase table has a row key with composite structure: customer_id (10 chars) + timestamp (10 digits). The access pattern is to query by customer_id and retrieve the latest entries. How should the DynamoDB table be designed to optimize performance?
Hard119A data engineer is reviewing an IAM policy that controls access to an S3 bucket. The policy is attached to a user group. The policy includes a condition that explicitly requires server-side encryption with SSE-S3 for all GetObject requests. The engineer notices that users are unable to download objects from the bucket. What is the likely cause?
Hard120A company has an Amazon Redshift cluster that stores petabytes of data. Queries are experiencing high disk usage due to large intermediate results. The data engineer needs to improve query performance without adding more nodes. Which action should the engineer take?
Hard121A company runs a real-time analytics platform using Amazon Kinesis Data Streams with a shard count of 10. The data is consumed by an AWS Lambda function that writes to an Amazon DynamoDB table. The DynamoDB table has a partition key of 'user_id' and a sort key of 'timestamp'. The table is provisioned with 5000 RCUs and 5000 WCUs. Recently, the application experienced increased write latency and throttling errors (ProvisionedThroughputExceededException) on the DynamoDB table. The CloudWatch metrics show that ConsumedWriteCapacityUnits averages 4500 with occasional spikes to 6000. The Lambda function’s concurrency is set to 1000. The data engineer suspects the issue is due to hot partitions. Upon investigation, the engineer finds that a small number of users generate a disproportionately large amount of data. Which course of action would best resolve the throttling while minimizing cost?
Hard122A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB in size. The migration must have minimal downtime and must be completed within a week. Which AWS service should the data engineer use to perform the migration?
Easy123A company runs a critical application on Amazon RDS for MySQL. To ensure high availability and automatic failover, the database is deployed as a Multi-AZ DB instance. The application uses read-heavy workloads. Which additional configuration should be used to offload read traffic without impacting write performance?
Medium124A data engineer manages an Amazon Redshift provisioned cluster that serves a nightly ELT workload. The cluster's largest fact table is loaded with new rows each night, and queries frequently filter on a date column and join to a customer dimension. The engineer wants to improve query performance and reduce the time spent vacuuming. Which TWO actions should the engineer take? (Choose two.)
Hard125Refer to the exhibit. A data engineer ran the CLI command to check the configuration of an RDS instance named 'mydb'. Which statement accurately describes the current configuration?
Hard126A data engineer needs to store a large volume of time-series data from IoT sensors. The data will be queried by timestamp and sensor ID, and the engineer wants to use a managed AWS database that can handle high write throughput and provide fast queries on recent data. The engineer also wants to automatically expire old data after 90 days to reduce storage costs. Which AWS service should the engineer use?
Easy127A data engineering team is designing a data lake on Amazon S3 for storing sensor data from IoT devices. The data is written in near real-time and needs to be queried using Amazon Athena. Which TWO configurations should the team implement to optimize query performance and minimize costs?
Medium128A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a Redshift table. The data is in Parquet format and is partitioned by date in S3. The engineer wants to load only the data for the last 7 days to reduce load time and cost. Which Redshift command should the engineer use?
Medium129An e-commerce application uses Amazon ElastiCache for Redis to cache product catalog data. The cache currently uses lazy loading. The team wants to ensure that frequently accessed product data is always fresh. Which caching strategy should they implement?
Easy130A company is using Amazon S3 to store sensitive customer data. The security team requires that all data be encrypted in transit and at rest. Additionally, they want to prevent any accidental public access. Which combination of actions should the data engineer take?
Hard131A data engineer is configuring an Amazon S3 bucket for a data lake. The bucket must store sensitive financial data and comply with a regulation that requires all data to be encrypted at rest with keys that are automatically rotated every year. The engineer also needs to audit key usage and control access to the keys separately from other AWS services. Which encryption option should the engineer choose?
Hard132A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it using Amazon Athena. The data is generated continuously and appended to S3 in small files. The engineer wants to optimize query performance and reduce costs. Which action should the engineer take?
Easy133A large e-commerce company uses Amazon DynamoDB to store shopping cart data. The table has a partition key of 'user_id' and a sort key of 'item_id'. The application performs frequent updates to the 'quantity' attribute for items in a user's cart. Recently, the operations team noticed that write requests are being throttled during peak shopping hours. The table is provisioned with 10,000 write capacity units (WCUs) and uses DynamoDB Accelerator (DAX) for read caching. The data engineer suspects that the throttling is due to hot partitions. The application uses a single AWS SDK client configured with retries. After reviewing the Amazon CloudWatch metrics, the engineer sees that the WriteThrottleEvents metric spikes for a few partition keys. The table has a high number of partitions. What should the data engineer do to resolve the throttling issue with minimal application changes?
Hard134A financial services company stores sensitive transaction data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys be automatically rotated every year. A data engineer needs to configure the S3 bucket to meet these requirements with minimal ongoing operational effort. Which solution should the engineer implement?
Hard135Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket. What is the effect of this policy?
Hard136A company uses Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed key that is rotated annually. Which encryption option should be used?
Easy137Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket named data-lake-bucket. The engineer wants to allow only GET requests from the corporate network (10.0.0.0/16) over HTTPS. However, users report that they cannot access objects even when connected to the corporate network. What is the issue?
Hard138A company is using Amazon RDS for MySQL and needs to automate backups with a retention period of 35 days. They also want to be able to restore to any point within the retention period. Which configuration should be used?
Medium139A company is designing a data store for IoT sensor data that is written once and never updated. The data must be stored with high durability and low cost. Which TWO AWS storage services are most suitable? (Choose TWO.)
Medium140A company needs to migrate an on-premises 10 TB PostgreSQL database to Amazon RDS for PostgreSQL with minimal downtime. Which AWS service should be used for the migration?
Easy141A company stores its application logs in Amazon S3. The logs are generated daily and need to be retained for 3 years for compliance. The logs are accessed frequently for the first 30 days, occasionally for the next 6 months, and rarely after that. The data engineering team wants to minimize storage costs while ensuring that logs are available for retrieval within 12 hours for the first 6 months and within 48 hours after that. The team also wants to automatically delete logs after 3 years. Which lifecycle policy should the team implement?
Easy142A data engineer is using Amazon Redshift and needs to improve query performance for a workload that involves frequent joins between a large fact table and a small dimension table. The dimension table is updated daily with new records. The engineer wants to minimize data movement during joins. Which Redshift distribution style should the engineer use for the dimension table?
Hard143A data engineer needs to store semi-structured JSON data from IoT devices. The data is written once, read rarely, but must be queryable using SQL. The storage cost must be minimized. Which storage solution should the engineer choose?
Medium144A data engineer is configuring an Amazon Redshift cluster and needs to optimize query performance for complex analytical queries that involve large joins. The engineer wants to reduce the amount of data movement during query execution. Which two actions should the engineer take? (Choose two.)
Medium145A company is using an Amazon RDS for PostgreSQL database to store application data. The data engineering team needs to run complex analytical queries that join multiple large tables. These queries are causing performance degradation on the production database. The team wants to offload the analytical workload to a separate system that can handle large-scale data processing. Which AWS service should the team use?
Easy146A company is using Amazon DynamoDB to store session data for a web application. The data engineer needs to ensure that the data is encrypted at rest. Which action should the data engineer take?
Medium147A data engineer is designing a data lake on Amazon S3 and needs to ensure that data is encrypted at rest. The company requires that encryption keys be managed by AWS and automatically rotated annually. The engineer also wants to audit key usage. Which S3 encryption option should the engineer choose?
Hard148A company stores application logs in Amazon S3 and needs to query them using standard SQL. The logs are in JSON format and are updated daily. The data engineering team wants a serverless solution that requires minimal management and can automatically discover the schema. Which AWS service should they use?
Easy149A data engineer needs to store transaction data that requires strong consistency, ACID transactions, and complex join queries. Which AWS service is most appropriate?
Easy150A data engineer applies the above IAM policy to a user. The user attempts to upload an object to the bucket 'my-data-lake' without specifying server-side encryption. What will happen?
Easy151A data engineer is troubleshooting an Amazon Redshift cluster that has been experiencing slow query performance. The engineer checks the system tables and finds that many queries are waiting on 'wlm_queued' time. The cluster has 10 nodes and uses automatic WLM. What is the most likely cause?
Hard152A data engineer is designing a data lake on Amazon S3 for a financial analytics workload. The raw data arrives as JSON files from an on-premises system. Analysts need to query the data using Amazon Athena with fast performance and minimal cost for queries that filter on a specific transaction date and customer ID. The engineer wants to convert the data to a columnar format that supports predicate pushdown and compression. Which storage format should the engineer choose?
Medium153A data engineer manages an Amazon Redshift cluster that experiences performance degradation during peak hours due to concurrent long-running queries and short ad-hoc queries competing for resources. The engineer wants to isolate the workloads so that short queries are not blocked by long-running ones, and to ensure that each workload gets a guaranteed share of memory and CPU. Which Redshift feature should the engineer implement?
Hard154A data engineer needs to migrate an on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which approach should they use?
Medium155A company uses DynamoDB with global tables in two AWS Regions. The data engineer observes that a write to the table in us-east-1 is not immediately visible in a read from eu-west-1. What is the most likely reason?
Hard156A company uses Amazon Redshift for analytics. They notice that some queries are slow due to data redistribution. The data engineer wants to minimize data movement across nodes. Which table design strategy should be used? (Choose TWO.)
Hard157Order the steps to set up an Amazon EMR cluster for processing data in S3 using Spark.
Medium158A company is migrating its on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is 2 TB and runs on a single server. Which AWS service should be used for the migration?
Medium159A data engineer is setting up an Amazon Redshift cluster and needs to load data from Amazon S3. The data is in CSV format and contains a large number of rows. The engineer wants to achieve the fastest possible load time. Which method should the engineer use?
Easy160A data engineer maintains an Amazon DynamoDB table that stores device telemetry. The table uses a partition key of deviceId and a sort key of timestamp, with on-demand capacity mode. A new fleet of devices writes data with a deviceId pattern that hashes to a small number of partitions, and the engineer observes throttling on writes even though consumed capacity is well below any configured limit. Which change addresses the root cause?
Hard161A data engineer is designing a data lake on Amazon S3 and needs to catalog data using the AWS Glue Data Catalog. The data is stored in Parquet format, partitioned by year/month/day. The engineer wants to query the data using Amazon Athena and ensure that partition pruning occurs to minimize query costs. Which action should the engineer take?
Hard162A data engineer runs the AWS CLI command to retrieve the lifecycle configuration of the 'my-data-lake' bucket. The output is shown in the exhibit. What is the effect of this lifecycle policy?
Medium163A data engineer is optimizing an Amazon S3 data lake that stores large volumes of JSON logs. The engineer wants to reduce storage costs and improve query performance in Amazon Athena. Which TWO actions should the engineer take? (Choose two.)
Medium164A data engineer is building an Amazon DynamoDB table that will store IoT sensor readings. Each reading has a deviceId (partition key) and a timestamp (sort key). The team wants to retrieve all readings for a device within the last 24 hours, and they also want to minimize the number of read capacity units consumed. Which access pattern should the engineer implement?
Hard165Refer to the exhibit. A data engineer runs the above CLI command and sees the output. The security team requires that the RDS instance not be accessible from the internet. Which change should the engineer make?
Medium166A data engineer is designing a data lake on Amazon S3 and needs to ensure that objects are automatically encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement achieves this?
Medium167A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. Each item has a partition key of deviceId and a sort key of timestamp. The table is configured with on-demand capacity mode. The engineer needs to retrieve all readings for a specific device within the last 24 hours, and the query must return results sorted by timestamp in ascending order. Which operation should the engineer use?
Medium168A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The engineer needs to retrieve all sessions for a specific user that started within the last 30 days. Which operation should the engineer use to achieve the lowest latency?
Hard169A company is using an Amazon RDS for MySQL database for an e-commerce application. During a sales event, the database experiences high read traffic, causing slow query performance. The company wants to reduce the read load on the primary database without changing the application code. Which solution meets these requirements?
Medium170A data engineer is setting up an Amazon S3 bucket to store large CSV files that will be queried using Amazon Athena. The engineer wants to minimize query costs and improve performance. The files are currently stored in a single prefix without any partitioning. The most common queries filter data by `year` and `month`. What should the engineer do to optimize the Athena queries?
Easy171A data engineer is troubleshooting a slow-running query on an Amazon Redshift cluster. The query involves joining two large tables. The engineer notices that the query plan shows a large number of distribution and broadcast operations. Which design change would most likely improve query performance?
Medium172Match each AWS Glue component to its role.
Medium173A data engineer is designing a data lake on Amazon S3 and needs to store data in a format that supports schema evolution and efficient columnar storage. The data will be queried using Amazon Athena and Amazon Redshift Spectrum. The engineer wants to minimize storage costs and improve query performance. Which storage format should the engineer choose?
Medium174A company uses Amazon S3 to store historical stock market data as CSV files. They run daily Amazon Athena queries to generate reports. Recently, the finance team reported that queries are timing out and costs have increased significantly. The data engineering team notices that the S3 bucket contains thousands of small files (average 100 KB) due to a misconfigured ingestion pipeline. They need to improve query performance and reduce costs without changing the existing reporting schedule. The team has access to AWS Glue and can create new tables. Which solution should they implement?
Medium175A company is migrating an on-premises Hadoop cluster to AWS. The cluster processes large files in CSV format using Apache Spark. Which data store should be used as the primary storage for the data lake to optimize cost and performance?
Medium176A financial services company stores transactional records in Amazon DynamoDB. Auditors require that any item be recoverable to its exact state from any point within the last 30 days, including after an accidental delete or overwrite caused by a faulty deployment. The table uses on-demand capacity and must remain highly available during recovery. Which feature should the data engineer enable?
Hard177Which THREE factors should a data engineer consider when choosing between Amazon RDS and Amazon DynamoDB for a new application? (Choose three.)
Hard178A data engineer needs to set up a new Amazon RDS for MySQL database for a web application. The application experiences variable read traffic and requires low read latency. The engineer needs to minimize downtime during maintenance and provide read scalability. Which configuration meets these requirements?
Easy179A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources and needs to be partitioned by year, month, day, and event type for efficient querying with Amazon Athena. Which S3 key prefix structure is most appropriate?
Easy180A data engineer is using AWS Lake Formation to manage fine-grained access control on an Amazon S3 data lake. The engineer has registered the S3 bucket as a Lake Formation data location and created a table in the AWS Glue Data Catalog. The engineer needs to grant a data analyst permission to query only specific columns (customer_id, order_date) in the sales table using Amazon Athena, while hiding other columns (credit_card_number, address). The analyst uses an IAM role that has no direct S3 permissions. Which action should the engineer take?
Hard181A data engineer is configuring an Amazon Redshift cluster for a reporting workload. The team needs to load data from Amazon S3 into a Redshift table and wants the fastest possible load while keeping the data compressed. Which approach should the engineer use?
Easy182A company stores customer transaction data in an Amazon DynamoDB table. The table has a partition key of CustomerID and a sort key of TransactionDate. The data engineering team needs to retrieve all transactions for a specific customer within a date range, and the queries must be efficient. Which DynamoDB operation should the team use?
Medium183A data engineer manages an Amazon S3 data lake with a bucket that has S3 Versioning enabled. A downstream analytics job accidentally overwrites thousands of current objects with corrupted data. The engineer must restore the previous good versions quickly and prevent the corrupted versions from being served. Which action should the engineer take?
Hard184A data engineer needs to store JSON documents that are accessed by a key-value pattern. The workload requires single-digit millisecond latency at any scale. Which AWS service is most appropriate?
Medium185A data engineer is designing a data lake on Amazon S3 that will be accessed by multiple AWS Glue ETL jobs. The engineer needs to ensure that the data is organized efficiently for querying and that sensitive columns are masked for certain users. Which TWO actions should the engineer take? (Choose TWO.)
Medium186A data engineer manages an Amazon S3 data lake with millions of small JSON files. To improve query performance with Amazon Athena, the engineer wants to compact these files into larger Parquet files. The engineer must also minimize ongoing storage costs. Which solution should the engineer implement?
Medium187A data engineer needs to store semi-structured JSON data that is accessed infrequently but must be retrievable within minutes. The data is generated by IoT devices and each object is about 500 KB. The engineer wants the most cost-effective storage solution. Which AWS service should be used?
Medium188A data engineer is managing an Amazon S3 data lake that contains raw JSON data. The engineer needs to optimize the data lake for query performance and cost when using Amazon Athena. The data is currently stored in a single S3 prefix without partitioning, and queries often filter on `event_type` and `event_date`. The engineer wants to implement best practices for Athena. Which TWO actions should the engineer take? (Choose two.)
Medium189A company stores application logs in Amazon S3 and uses AWS Glue crawlers to populate the AWS Glue Data Catalog. A data engineer needs to query the logs with Amazon Athena. The logs are partitioned by year/month/day in S3, but Athena queries are scanning all partitions and returning errors about missing partitions. What should the engineer do to enable partition pruning?
Easy190A media company stores millions of small JSON files in an Amazon S3 bucket and queries them with Amazon Athena. Analysts report that queries scan far more data than expected, and costs are rising. The data engineer confirms that the files are uncompressed, use no partitioning, and are stored as newline-delimited JSON. Which change will MOST reduce the data scanned per query?
Hard191A data engineer is designing a data lake on Amazon S3 for a retail company. The company ingests point-of-sale data as small JSON files every few minutes, totaling about 5 GB per day. Analysts query the data with Amazon Athena, and costs are rising due to many small files and full scans. The engineer wants to reduce Athena query costs and improve performance while keeping the data in S3. Which TWO actions should the engineer take? (Choose two.)
Hard192A data engineer needs to migrate an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB and has a continuous stream of write operations. The migration should minimize downtime. Which AWS service should be used?
Medium193A company runs an Amazon RDS for PostgreSQL instance for an OLTP application. The database size is 500 GB. The company wants to minimize downtime during backups and ensure point-in-time recovery (PITR) for the last 7 days. Which TWO features should the company use? (Choose TWO.)
Hard194A company stores sensitive data in an Amazon S3 bucket. A compliance requirement mandates that all data must be encrypted at rest with a key that is automatically rotated every year. The company also needs to maintain an audit trail of who used the key. Which solution meets these requirements?
Medium195A data engineer is designing a solution to ingest streaming data from Amazon Kinesis Data Streams into an Amazon Redshift cluster for near-real-time analytics. The engineer needs to ensure that data is loaded efficiently and that the Redshift cluster can handle the ingestion load without impacting query performance. Which approach should the engineer use?
Medium196A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it with Amazon Athena. The engineer wants to minimize query costs and improve performance. Which action should be taken?
Easy197A data engineer is troubleshooting slow query performance on an Amazon Redshift cluster. The cluster has 10 nodes and is using automatic distribution style. The engineer suspects that data distribution is causing excessive data movement. Which steps should the engineer take to diagnose and resolve the issue? (Choose THREE.)
Hard198An application uses the 'orders' DynamoDB table with the schema and provisioned throughput shown in the exhibit. The application frequently queries by customer_id (range key) without specifying the order_id (partition key). What is the most likely impact on performance?
Hard199A company is using Amazon RDS for MySQL with Multi-AZ deployment. The database size is 2 TB and the workload is read-heavy. To improve read performance, which option should be used?
Medium200A data engineer is using AWS Glue to catalog data stored in Amazon S3. The engineer needs to run an AWS Glue ETL job that reads from a large dataset in Parquet format, performs transformations, and writes the output to Amazon Redshift. The job must handle data skew and optimize performance. Which AWS Glue feature should the engineer use to address data skew during the join operation?
Hard201A data engineer is troubleshooting a slow-running query on Amazon Redshift. The query scans a large table but returns few rows. Which diagnostic step should be taken first?
Medium202A company is migrating a large Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime and preserve data consistency. Which THREE AWS services or features should be used?
Hard203A data engineer is using Amazon Redshift to store sales data. The engineer needs to ensure that the data is encrypted at rest and that encryption keys are managed by AWS. The engineer also wants to minimize administrative overhead. Which Redshift encryption option should the engineer use?
Easy204A company is using Amazon RDS for PostgreSQL with Multi-AZ deployment. The primary instance fails and a failover occurs. After the failover, the application cannot connect to the database. What is the MOST likely cause?
Medium205A company is storing sensitive user data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using a customer-managed key stored in AWS KMS. The bucket policy must deny any PUT request that does not include the appropriate encryption header. Which bucket policy condition key should be used?
Medium206A company is designing a data lake on Amazon S3 for analytics. The data includes sensitive personally identifiable information (PII). Which TWO actions should the company take to protect the data? (Choose TWO.)
Medium207A data engineer runs the above CLI command to describe the DynamoDB table 'Orders'. The table has a partition key 'OrderID' and sort key 'CustomerID'. Which query operation is most efficient for retrieving all orders for a specific customer?
Medium208Refer to the exhibit. A data engineer applies this bucket policy to an S3 bucket. A user within the 10.0.0.0/24 IP range attempts to upload an object to the bucket using an HTTP (non-HTTPS) request. What is the outcome?
Hard209A data engineer maintains an Amazon DynamoDB table that stores IoT telemetry. The table uses a partition key of deviceId and a sort key of timestamp, with on-demand capacity. A few very active devices generate millions of writes per hour while thousands of other devices write sporadically. The engineer observes throttling on writes for the active devices and wants to reduce it with the least application change. Which action should the engineer take?
Hard210A data engineer is migrating an on-premises MongoDB database to Amazon DocumentDB. Which migration strategy minimizes downtime?
Medium211A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Apache Parquet format and partitioned by year, month, and day. The engineer notices that AWS Glue crawlers are taking a long time to run and are not correctly identifying new partitions. The engineer needs to improve the crawler performance and ensure new partitions are added automatically. Which action should the engineer take?
Hard212A financial services company stores transaction data in Amazon RDS for PostgreSQL. The company requires that all changes to the database be logged for audit purposes, including before and after images of updated rows. Which feature should the data engineer enable?
Hard213A data engineer needs to store archival data that is rarely accessed but must be retained for 7 years. The data should be retrievable within 12 hours. Which Amazon S3 storage class is MOST cost-effective?
Easy214A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data should be stored cost-effectively. Which storage solution meets these requirements?
Easy215A company stores sensitive data in Amazon S3. They need to ensure that all objects are encrypted at rest. Which approach meets this requirement with minimal effort?
Medium216A data engineer needs to create a table in Amazon Athena that reads JSON data stored in Amazon S3. The JSON records are stored in a single file, one JSON object per line. The engineer wants Athena to automatically discover the schema and create the table without manually defining columns. Which AWS service or feature should the engineer use?
Medium217A data engineer is designing a data lake on Amazon S3. The data includes customer PII that must be encrypted at rest. The company also requires that the encryption keys be rotated automatically every year. Which encryption solution should the engineer use?
Easy218A company uses Amazon DynamoDB with provisioned capacity. During a sales event, write traffic spikes and some requests receive ProvisionedThroughputExceeded exceptions. The reads are within limits. The data engineer needs to minimize latency for the spike without manual intervention. Which solution is MOST cost-effective?
Hard219A company stores critical financial data in Amazon DynamoDB. To meet compliance requirements, the data must be encrypted at rest with a customer-managed key. Which solution should the data engineer implement?
Easy220A data engineer is designing a data lake on Amazon S3 using AWS Lake Formation. The engineer needs to grant fine-grained access to specific columns and rows of a table to different analysts. Which two actions should the engineer take to meet these requirements? (Choose two.)
Hard221A data engineer is using Amazon Redshift and needs to improve the performance of complex queries that join large tables. The engineer has already set the distribution style to KEY on the join columns. What additional step should the engineer take to optimize the join performance?
Medium222A company is using Amazon DynamoDB with on-demand capacity for a gaming application. During a new game launch, write traffic spikes to 50,000 writes per second, but the application experiences throttling. The DynamoDB table has a partition key of 'game_id' and a sort key of 'timestamp'. What is the MOST likely cause of throttling?
Hard223A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for ad-hoc querying using SQL. Which AWS service should be used?
Medium224A company is using Amazon S3 as a data lake. The data engineer needs to ensure that all objects uploaded to a specific bucket are automatically replicated to a bucket in another AWS Region for disaster recovery. Which configuration should the engineer implement?
Easy225A company is using Amazon DynamoDB with auto scaling enabled. During a marketing campaign, write traffic spikes, and some write requests fail with ProvisionedThroughputExceededException. The auto scaling policy has a target utilization of 70% and a maximum capacity that is high enough. What is the most likely cause of the throttling?
Hard226A company is designing a data lake on Amazon S3. The data includes personal identifiable information (PII). The data engineer must ensure that only authorized users can access the data, and that access is logged for auditing. Which combination of services should the data engineer use?
Hard227A data engineer is troubleshooting an access denied error when an AWS Lambda function tries to decrypt an object encrypted with the KMS key 'abc123'. The Lambda function's execution role has the above policy attached. What is the likely cause of the error?
Hard228A company uses Amazon DynamoDB as the primary data store for a gaming application. The application stores user profiles and game state. During peak hours, the application experiences throttling on writes to the UserProfiles table. The table's read capacity is underutilized. Which solution should resolve the write throttling?
Easy229A data engineer is designing a data lake on Amazon S3. The data consists of sensitive personally identifiable information (PII) that must be encrypted at rest. The company requires that encryption keys be rotated every 90 days and that access to the keys be logged. Which encryption solution meets these requirements?
Medium230A data engineer is building a data lake on Amazon S3 and must choose the optimal file format for a dataset that is queried by Amazon Athena. The queries typically select a few columns from wide tables containing hundreds of columns, and the data volume is in terabytes. The engineer wants to minimize query scan costs and improve performance. Which file format should the engineer use?
Medium231A CloudFormation template includes this IAM policy for a cross-account S3 upload use case. What is the purpose of the condition?
Hard232A data engineer runs the above command and gets the output. What does the 'MFADelete' setting imply?
Medium233A company runs an Amazon RDS for PostgreSQL instance that stores financial data. The company requires point-in-time recovery (PITR) with a retention period of 35 days. Additionally, the company needs to create a new database from a specific snapshot every night for testing. Which combination of actions should the data engineer take to meet these requirements?
Hard234A data engineer needs to store log files from multiple applications in a centralized location. The logs are generated in JSON format and each log entry is about 1 KB. The engineer needs to query the logs occasionally using SQL-like queries. Which AWS service is most appropriate?
Easy235A company is using Amazon EMR to process large datasets stored in Amazon S3. The data engineer wants to reduce the time it takes to read data from S3 by optimizing the data format. Which file format should the engineer recommend?
Easy236Which TWO are benefits of using Amazon S3 Object Lock? (Choose TWO.)
Hard237A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak hours. The cluster uses a single node type and has no concurrency scaling enabled. Analysis shows that many long-running queries are queued behind short ad-hoc queries, causing delays for critical reports. The engineer needs to ensure that critical reports run promptly without affecting ad-hoc queries. Which solution meets these requirements?
Hard238A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a table. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition and ensure that the load is efficient and cost-effective. Which method should the engineer use?
Hard239A data engineer needs to set up a new Amazon RDS for PostgreSQL database for a production workload. The database must be highly available and resilient to a single Availability Zone failure. Which configuration should the engineer choose?
Hard240A company is using Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key. The data engineer must ensure that only a specific IAM role can decrypt the data. Which policy should the data engineer attach to the KMS key?
Hard241A data engineer is configuring an Amazon S3 bucket that will receive raw clickstream files from a mobile application. The engineer must ensure that the objects are protected against accidental overwrites and deletions for a defined retention period, and that the protection cannot be removed or shortened by any user, including the account root user. Which S3 feature should the engineer use?
Easy242Refer to the exhibit. A data engineer runs the above AWS CLI command to view the table metadata in the AWS Glue Data Catalog. The data is stored as CSV in S3 with partitions by year and month. When querying the table using Amazon Athena, no data is returned. What is the most likely cause?
Medium243Refer to the exhibit. A data engineer notices that the Redshift cluster 'mycluster' does not have automated backups beyond 7 days. However, the compliance team requires a minimum of 35 days of backup retention. What should the engineer do?
Medium244A data engineer needs to store streaming data from IoT devices for real-time analytics. The data has a fixed schema and requires low-latency queries. Which AWS service should be used?
Easy245A data engineer is designing a data lake on Amazon S3. The data is frequently accessed by multiple analytics services, and the company needs to enforce fine-grained access control based on data tags. Which combination of AWS services should be used?
Hard246A data engineer needs to store JSON documents that are frequently accessed by a low-latency web application. The data does not require complex queries, and the access pattern is primarily by a key. Which AWS service is most appropriate?
Easy247A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a folder structure. The engineer wants to use AWS Glue crawlers to automatically infer schemas and create tables in the AWS Glue Data Catalog. The crawler should run daily to detect new files and schema changes. Which configuration should the engineer use for the crawler?
Easy248A company uses Amazon DynamoDB with on-demand capacity for a gaming application that experiences unpredictable traffic spikes. The application reads the same set of 'hot' items frequently. Users report high latency during peak hours. Which action would MOST effectively reduce read latency for the hot items?
Hard249The exhibit shows an S3 bucket policy. What is the effect of this policy?
Medium250A data engineer is building a data lake on Amazon S3. The engineer needs to catalog metadata for data stored in Parquet format and make it queryable by Amazon Athena. The data is partitioned by year, month, and day in the S3 path. Which AWS service should the engineer use to create and manage the table definitions and partitions?
Easy251A company stores time-series sensor data in Amazon S3. They need to query the data using SQL with minimal latency and no infrastructure management. Which service should they use?
Easy252A data engineer is designing an Amazon S3 data lake and needs to enforce schema-on-read for a dataset that is queried by Amazon Athena. The data is stored as Parquet files partitioned by year, month, and day. The engineer wants to minimize the amount of data scanned by queries that filter on a specific date range. Which approach should the engineer take?
Medium253A data engineer needs to store semi-structured JSON transaction logs for analytics. The logs are written once and rarely accessed. The storage must be cost-effective. Which AWS service should be used?
Easy254A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources in Parquet format, and the schema evolves over time. Which approach allows querying the data with Amazon Athena while supporting schema evolution?
Medium255A company runs an Apache Spark job on Amazon EMR that writes output to an S3 bucket. The job fails with the error 'S3AccessDeniedException' when writing the final output, but earlier stages succeed. The EMR cluster uses a service role and an instance profile. The S3 bucket policy allows access from the VPC only. What is the MOST likely cause?
Hard256A company has an Amazon RDS for MySQL DB instance with read replicas. The primary DB instance fails. What is the correct procedure to promote a read replica to become the new primary?
Medium257A data engineering team is using Amazon EMR to process large datasets stored in Amazon S3. The cluster uses Spot Instances for cost savings. During processing, the team notices that tasks are failing due to Spot Instance interruptions. The team needs to make the EMR job resilient to Spot interruptions without increasing costs significantly. Which solution should they implement?
Medium258A company is migrating an on-premises Apache Cassandra database to Amazon Keyspaces. The database has a table with a partition key of 'user_id' and a clustering column of 'timestamp'. The application frequently queries the last 10 records for a given user. Which table design in Keyspaces would provide the BEST query performance for this access pattern?
Medium259A data engineer sees this AWS Glue table definition in the Data Catalog. The engineer wants to query this table with Amazon Athena, but the query returns zero rows. What is the MOST likely cause?
Medium260A company uses Amazon S3 to store sensitive data. The security team wants to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted at rest using server-side encryption with AWS KMS managed keys (SSE-KMS). Which bucket policy statement should be added to enforce this requirement?
Medium261A company is using Amazon DynamoDB for a gaming application. They want to store player session data that expires after 24 hours. Which DynamoDB feature should be used?
Easy262Which TWO actions can reduce the cost of an Amazon S3 bucket that stores infrequently accessed data? (Choose 2.)
Medium263A data engineer is setting up Amazon S3 event notifications to trigger an AWS Lambda function when new objects are uploaded. Which TWO actions are required to enable this?
Easy264A data engineer must choose a storage service for a new application that requires single-digit millisecond latency at any scale, a flexible schema, and automatic scaling of throughput without provisioning capacity. The access pattern is key-value lookups by user ID with occasional range queries on a sort key. Which AWS service should the engineer select?
Easy265A data engineer manages an Amazon Redshift cluster that experiences performance degradation during complex analytical queries. The engineer notices that some queries spill to disk. The engineer wants to improve query performance by optimizing the distribution style and sort keys. Which action should the engineer take first?
Hard266A data engineer manages an Amazon Redshift cluster that runs a nightly ETL load followed by complex analytical queries. Users report that queries during the day are slower than expected, and the team wants to isolate the ETL workload so it cannot consume resources needed by the analytical queries. The cluster uses provisioned nodes. What is the MOST appropriate solution?
Medium267A media company stores video files in an S3 bucket. The files are processed by a fleet of EC2 instances that read the files, add watermarks, and write the output back to the same bucket. Recently, the processing jobs have been failing with '500 Internal Server Error' and '503 Slow Down' errors. The data engineer checks the S3 bucket metrics and sees that the PUT/GET request rate is consistently above 5,500 requests per second for a single prefix. The engineer needs to resolve the errors with minimal changes to the application code. Which course of action should the engineer take?
Medium268A data engineer is building a data lake on Amazon S3 and must enforce that all objects containing personally identifiable information are encrypted with a customer managed AWS KMS key, while allowing automatic key rotation and audit of key usage. Objects must remain readable by an AWS Glue job and an Amazon Athena workgroup. Which configuration should the engineer choose?
Medium269A company uses Amazon Redshift for data warehousing. The data engineering team notices that queries are slow due to high disk I/O. The team wants to improve query performance without changing the cluster configuration. Which action should the team take?
Medium270A company uses Amazon S3 to store large datasets. The data engineering team needs to provide access to specific objects in the bucket to external partners using presigned URLs. Each URL should expire after 12 hours. The team wants to ensure that the presigned URLs cannot be used to access other objects in the bucket. Which approach should be taken?
Hard271A data engineer needs to store JSON documents that are accessed by a serverless application using AWS Lambda. The documents are frequently updated and need low latency (single-digit milliseconds) for read and write operations. Which AWS service should the engineer use?
Easy272A company uses Amazon Redshift for analytics. The data engineering team wants to improve query performance for frequently used aggregate queries. Which TWO actions would help achieve this?
Medium273A data engineer is using Amazon Redshift and needs to improve query performance for a large fact table that is frequently joined with a much smaller dimension table. The engineer wants to minimize data movement during joins. Which distribution style should be used for the dimension table?
Medium274A data engineer notices that an Amazon Redshift cluster’s storage usage is increasing rapidly due to many UPDATE and DELETE operations. The engineer needs to reclaim storage space and improve query performance. Which action should be taken?
Hard275A data engineer needs to store semi-structured data (JSON logs) from thousands of IoT devices. The data must be schema-less, highly scalable, and support low-latency queries by device ID and timestamp. Which AWS service should the engineer use?
Easy276A company stores application logs in an Amazon S3 bucket. A compliance policy states that log objects must be retained for exactly 90 days and then permanently deleted, and that no one, including administrators, should be able to delete them earlier. The data engineer must enforce this with the least effort. What should the engineer do?
Easy277A data engineer is configuring an Amazon S3 bucket to store sensitive financial data. The company requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption key be automatically rotated every year. The engineer creates a KMS customer managed key and enables automatic rotation. When uploading objects using the AWS CLI, the engineer uses the --sse aws:kms parameter but does not specify a key ID. What is the result of this configuration?
Hard278Which THREE factors should be considered when choosing a partition key for an Amazon DynamoDB table?
Hard279A data engineer manages an Amazon S3 data lake with millions of small JSON files ingested continuously. Amazon Athena queries over this data are slow and expensive because each query scans many small objects. The engineer wants to improve query performance and reduce cost without changing the data content. Which solution should the engineer implement?
Medium280A data engineer needs to store semi-structured JSON data from IoT devices. The data is written frequently and read occasionally. Which AWS service is MOST cost-effective for this use case?
Easy281A data engineer is using AWS Glue to process data stored in Amazon S3. The engineer needs to ensure that the AWS Glue job can access the S3 bucket securely without hardcoding credentials. Which approach should the engineer use?
Medium282A data engineer is designing a multi-region disaster recovery solution for Amazon RDS for PostgreSQL. The primary region must have a standby in a different Availability Zone, and the secondary region must have a readable replica that can be promoted in case of failure. Which configuration meets these requirements?
Hard283A company uses Amazon DynamoDB for a gaming application. The application experiences throttling during peak hours. The table's read and write capacity is provisioned. Which TWO actions can reduce throttling?
Medium284A media company stores millions of thumbnail images in an Amazon S3 bucket. Analysts run ad hoc queries against the image metadata, which is kept as JSON objects in the same bucket. Query latency is unpredictable and costs are rising because Athena scans large volumes of JSON for every query. The team wants faster queries and lower scan cost while keeping the data in S3 and queryable with SQL. Which change should the data engineer make?
Medium285A data engineer is deploying an Amazon Redshift cluster that must be accessible only from within a private VPC and must not have a public IP address. The cluster will be queried by an Amazon EMR cluster in the same VPC and by on-premises BI tools over a VPN connection. Which configuration should the engineer choose?
Medium286A data engineer is configuring an Amazon S3 bucket that stores sensitive customer records for analytics. The security team requires that all data be encrypted at rest with keys that are rotated automatically every year and that access be auditable per key. The engineer must minimize operational overhead. Which encryption configuration should be used?
Medium287A company uses Amazon DynamoDB to store user session data. The table has a partition key of user_id and a sort key of session_start. The workload is read-heavy and eventually consistent reads are acceptable. The table is provisioned with 1000 RCUs and 500 WCUs. During peak hours, the application experiences throttling on read operations, but CloudWatch shows that the consumed read capacity is well below the provisioned amount. What is the most likely cause of the throttling?
Hard288A data engineer needs to store JSON documents that are frequently read and written by a web application. The data has a flexible schema and requires low-latency queries on primary key lookups. Which AWS service is MOST suitable?
Easy289A data engineer needs to store and analyze time-series data from IoT devices. The data volume is 10 GB per day, and the queries are mostly on the most recent 7 days of data. The engineer wants to minimize storage costs while retaining historical data for 1 year. Which combination of AWS services is most cost-effective?
Medium290A company is using Amazon S3 for data lake storage. They need to query the data directly using SQL without loading it into a database. Which AWS service should be used?
Easy291A company wants to use Amazon Redshift Spectrum to query data in Amazon S3. The data is in Parquet format and partitioned by date. Which step is required to enable Redshift Spectrum?
Easy292A company needs to store files that are accessed by multiple EC2 instances in a VPC. The files must be concurrently accessible and durable. Which storage solution should the data engineer choose?
Easy293A data engineer manages an Amazon DynamoDB table for order events. Reads and writes are evenly spread across a partition key with very high cardinality, but during flash sales the table throttles with ProvisionedThroughputExceededException even though consumed capacity is below the provisioned total. Which cause is MOST likely?
Hard294A data engineering team is using AWS Glue to catalog data in an S3 data lake. They have a Glue crawler that runs daily to update the Data Catalog. Recently, they noticed that the crawler is taking longer to run and sometimes fails because of a timeout. The team suspects the issue is due to the large number of small files in the S3 bucket. They need to improve crawler performance and reliability. Which solution should they implement?
Easy295A data engineer is setting up Amazon S3 bucket policies for a data lake. Which TWO statements are true regarding S3 bucket policies? (Choose TWO.)
Easy296A company uses Amazon S3 as its data lake. A data engineer needs to enforce encryption of data at rest using server-side encryption with AWS KMS. Which S3 bucket property should be configured?
Easy297Refer to the exhibit. A data engineer creates an Amazon Redshift table with the above DDL. The engineer runs a query to find all orders for a specific customer within a date range. Which statement about query performance is correct?
Easy298A data engineer is optimizing an Amazon RDS for MySQL database that experiences high write throughput. The engineer wants to improve write performance and reduce latency. Which TWO database-level configuration changes can help achieve this?
Medium299A company has an Amazon DynamoDB table with a provisioned write capacity of 1000 WCU. During a flash sale, the write traffic spikes to 5000 WCU for 10 minutes. The table is not auto-scaled. Which action should the data engineer take to handle the spike without throttling?
Hard300A data engineer is migrating a large Oracle data warehouse to Amazon Redshift. The engineer needs to ensure optimal performance. Which TWO practices should the engineer follow?
Medium301A retail company uses Amazon DynamoDB to store product catalog data. The table has a partition key of ProductID and a sort key of Category. The company needs to retrieve all products in a specific category, sorted by ProductID. Which operation should be used?
Easy302Which THREE storage classes in Amazon S3 are designed for infrequently accessed data with millisecond retrieval times? (Select THREE.)
Medium303A company is using Amazon RDS for MySQL with Multi-AZ deployment. The primary DB instance experiences a hardware failure, causing automatic failover to the standby. After the failover, the application reports that the database endpoint is unreachable for about 60 seconds. What is the MOST likely cause?
Medium304A data engineer needs to store large amounts of data that is accessed infrequently but must be retrieved immediately when needed. Which Amazon S3 storage class is most cost-effective?
Easy305A company uses Amazon RDS for PostgreSQL to store customer data. The data engineer needs to ensure that the database can be restored to any point in time within the last 35 days. The engineer also wants to minimize the impact on the production database during backups. What should the engineer do?
Easy306A company uses Amazon DynamoDB with global tables in three AWS Regions. The data engineer needs to ensure that writes to the table in us-east-1 are replicated to other regions with minimal latency. Which DynamoDB feature should be used?
Medium307A company has a DynamoDB table with a partition key of 'user_id' and a sort key of 'timestamp'. They need to query all items for a user within a date range. Which query operation should be used?
Hard308A company runs an Amazon Redshift cluster with 10 RA3 nodes. The data warehouse stores 50 TB of data. The company notices that queries are slow and the cluster's storage utilization is high. The data engineer needs to improve query performance and reduce storage costs without changing the cluster's node count. Which action should the engineer take?
Hard309A company uses an Amazon RDS for MySQL DB instance with Multi-AZ deployment. The primary DB instance fails unexpectedly. What happens to the database endpoint?
Easy310Refer to the exhibit. A data engineer configured the lifecycle policy shown. The 'logs/' prefix contains important audit logs. After 365 days, what happens to the objects?
Medium311Which TWO statements about Amazon Redshift data distribution are correct? (Choose two.)
Easy312A company stores application logs in Amazon S3 in JSON format. The logs are partitioned by year/month/day. A data engineer needs to create a table in the AWS Glue Data Catalog so that Amazon Athena can query the logs efficiently. The engineer wants to minimize query costs and ensure that new partitions are automatically recognized. Which combination of actions should the engineer take?
Medium313A data engineer manages an Amazon DynamoDB table used for a high-traffic gaming leaderboard. The table uses on-demand capacity mode and has a partition key of UserId (string) with no sort key. The leaderboard must retrieve the top 100 scores across all users. Currently, the engineer scans the entire table and sorts the results in application code, which takes several seconds and consumes large amounts of read capacity. What should the engineer do to improve the performance of retrieving the top scores?
Medium314A company uses Amazon DynamoDB for a gaming leaderboard. The table has a partition key of 'GameId' and a sort key of 'Score'. The application needs to query the top 10 scores for a given game. Which DynamoDB feature should be used for optimal performance?
Hard315A company uses Amazon RDS for MySQL with Multi-AZ deployment. The primary instance fails, and automatic failover occurs. After failover, the application experiences higher latency. What is the most likely cause?
Medium316A company needs to store relational data that requires complex joins and transactional consistency. The workload is predictable and the data size is less than 500 GB. Which AWS service is MOST cost-effective for this use case?
Easy317A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The data engineer needs to ensure minimal downtime during migration. Which AWS service should be used to facilitate the migration?
Easy318A data engineer is designing a data lake on Amazon S3. The data lake will store raw data, transformed data, and curated datasets. The engineer needs to ensure that raw data is immutable (never overwritten or deleted) and that only authorized users can access the transformed data. Which combination of S3 features should the engineer use?
Easy319A data engineer is designing a data warehouse on Amazon Redshift. The workload includes many ad-hoc queries that filter on a high-cardinality column, such as customer_id, and join large dimension tables. The engineer wants to improve query performance by choosing an appropriate distribution style and sort key. Which combination should the engineer use?
Hard320A data engineer is configuring an Amazon Redshift cluster for a workload that runs large nightly ELT jobs loading data from Amazon S3 and then executes complex analytical queries. The team wants to improve query performance and reduce the time the cluster spends on data loading. Which TWO configuration choices should the engineer make? (Choose two.)
Medium321A data engineer is setting up an Amazon DynamoDB table to store user session data. The table must handle sudden spikes in read traffic during peak hours, and the engineer wants to minimize operational overhead while ensuring consistent performance. The table's read capacity mode should automatically adjust to traffic changes. Which capacity mode should the engineer choose?
Easy322Which TWO statements are true about Amazon Redshift distribution styles? (Choose TWO.)
Medium323A data engineer is configuring an AWS Glue ETL job to read from an Amazon S3 bucket that contains Apache Parquet files partitioned by year, month, and day. The engineer wants the job to only process data for the year 2023 and month 10, and to minimize the amount of data scanned. The Glue job uses the Glue Data Catalog table `sales_data` with the correct partition structure. What is the MOST efficient way to configure the job to read only the required partitions?
Medium324A data engineer needs to store JSON documents that are frequently updated and require ACID transactions. Which AWS database service is most appropriate?
Easy325A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB and has a 24/7 uptime requirement. The migration must minimize downtime. Which approach should be used?
Medium326A data engineer applies the following IAM policy to an IAM user: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::example-bucket/*", "Condition": { "StringEquals": { "s3:x-amz-server-side-encryption": "AES256" } } } ] } ``` The user attempts to download an object from the bucket 'example-bucket' that is encrypted with SSE-S3 (AES256). Will the request succeed?
Medium327A company stores sensitive data in Amazon S3 and needs to ensure that data is encrypted at rest. The security team requires that the company manage its own encryption keys and have the ability to audit key usage. Which S3 encryption option should the data engineer choose?
Easy328A company needs to store JSON documents that are frequently read and written by a web application. The data must be highly available and durable across multiple Availability Zones. Which AWS database service meets these requirements?
Easy329A company runs a MySQL database on Amazon RDS. The database size is 500 GB and is experiencing high read traffic. The team wants to improve read performance with minimal operational overhead. Which action should they take?
Easy330A data engineer notices that an Amazon Redshift cluster is running low on disk space. The cluster has three nodes of type dc2.large. Which action will increase the available storage capacity?
Medium331A data engineer is designing a data lake on Amazon S3. Data is ingested from multiple sources in JSON format. The engineer needs to optimize query performance for Amazon Athena while minimizing storage costs. Which storage strategy should the engineer use?
Medium332Which TWO AWS services can be used to automatically back up an Amazon RDS for SQL Server DB instance? (Choose TWO.)
Easy333A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by year, month, and day. The engineer needs to ensure that AWS Glue crawlers correctly identify the partitions and that Amazon Athena queries can efficiently prune partitions. Which action should the engineer take?
Medium334A company is using Amazon S3 to store critical data and needs to ensure that objects are automatically transitioned to S3 Glacier Deep Archive after 180 days to reduce costs. Which S3 lifecycle action should be configured?
Easy335A company needs to store archival logs that must be retained for 10 years. The logs are accessed infrequently, but when accessed, retrieval must occur within 12 hours. Which storage class is MOST cost-effective?
Easy336A data engineer manages an Amazon S3 data lake that holds sensitive customer transaction logs. Compliance requires that all objects be encrypted at rest with keys that the company rotates every 90 days and fully controls, including the ability to immediately revoke access and audit key usage separately from other AWS accounts. The engineer must choose an encryption method that meets these requirements with minimal operational overhead. Which solution should the engineer implement?
Medium337A data engineer is using Amazon Athena to query data stored in an S3 bucket. The queries are running slowly. Which THREE actions can improve query performance?
Hard338A data engineer is designing a data lake on Amazon S3. Which feature should be used to manage the lifecycle of objects and move them to cheaper storage classes automatically?
Easy339Refer to the exhibit. A data engineer needs to connect to the Redshift cluster from an EC2 instance in the same VPC. The engineer can ping the EC2 instance but cannot connect to Redshift using the endpoint address and port 5439. What is the most likely cause?
Medium340A data engineer is designing a real-time analytics pipeline that ingests clickstream data into Amazon Kinesis Data Streams. The data must be stored in Amazon S3 for later analysis with Amazon Athena. The engineer needs the data to be queryable with minimal latency and wants to avoid managing complex ETL jobs. Which solution should the engineer use?
Hard341A data engineer is building a data lake on Amazon S3. The engineer needs to store structured data that will be queried by Amazon Athena. The data is currently in CSV format and is partitioned by date. The engineer wants to improve query performance and reduce the amount of data scanned. Which action should the engineer take?
Easy342A data engineer needs to store semi-structured JSON log files from multiple sources and query them using SQL. The data is rarely updated and access frequency is low. Which storage solution is MOST cost-effective?
Easy343A company is migrating its on-premises MySQL database to Amazon RDS for MySQL. They want to minimize downtime and ensure data consistency. Which AWS service should be used for the migration?
Easy344Which TWO actions can help improve the read performance of an Amazon DynamoDB table that is experiencing throttling? (Choose two.)
Medium345A data engineer runs the above SQL commands on an Amazon Redshift cluster. The table 'users' is created with DISTSTYLE EVEN. What is the effect of the DISTSTYLE EVEN on query performance?
Easy346A company uses AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by date. The company wants to improve query performance in Amazon Athena and reduce costs. Which THREE actions should the company take? (Choose THREE.)
Medium347A company wants to enforce that all data in an S3 bucket is encrypted at rest using AWS KMS. Which bucket policy condition key should be used?
Easy348Match each AWS monitoring tool to its primary use.
Medium349A data engineer stores Apache Parquet files in an Amazon S3 data lake partitioned by dt=YYYY-MM-DD. Analysts query the data with Amazon Athena, and monthly reports that scan one month of data are slow and expensive. The engineer confirms that queries filter on the dt column. Which action will MOST effectively reduce the amount of data scanned by these reports?
Medium350A logistics company stores shipment tracking events in an Amazon DynamoDB table. The table uses a partition key of shipment_id and a sort key of event_timestamp. Analysts frequently run queries that filter by shipment_id and a range of event_timestamp values. The data engineer must ensure these queries are efficient and consume minimal read capacity. What should the data engineer do?
Easy351A data engineer needs to store time-series data from IoT devices. The data is write-heavy and requires low-latency queries by device ID and timestamp. The data volume is expected to grow to terabytes. Which AWS database service is most suitable?
Easy352A company has an Amazon RDS for MySQL database that is experiencing performance issues due to a large number of read requests. The application is read-heavy and can tolerate eventually consistent reads. Which action will reduce the load on the primary database with the least operational overhead?
Hard353A data engineer is managing an Amazon S3 data lake that contains millions of small files. The engineer needs to optimize query performance in Amazon Athena and reduce costs. The data is stored in Parquet format and is partitioned by date. Which action should the engineer take to improve performance and reduce costs?
Hard354A company is migrating a legacy data warehouse to Amazon Redshift. They need to choose a distribution style to minimize data movement during joins. Which THREE factors should they consider?
Hard355A company uses Amazon Redshift for its data warehouse. The data engineering team loads data daily from Amazon S3 using COPY commands. Recently, the load performance has degraded because the S3 bucket contains many small files. The team needs to optimize the COPY operation to improve performance. Which approach should they take?
Easy356A data engineer is designing a DynamoDB table for an application that requires strongly consistent reads and supports a global secondary index (GSI). The engineer needs to ensure that queries on the GSI return the most up-to-date data. Which statement about DynamoDB read consistency is correct?
Hard357A data engineer is designing a data lake on Amazon S3 that will store sensitive financial data. The engineer needs to implement encryption at rest and ensure that only authorized users can access the data. Which TWO actions should the engineer take to meet these requirements? (Choose TWO.)
Medium358A data engineer is troubleshooting a slow Amazon Redshift query that joins several large tables. The query plan shows a large number of broadcasts. Which design change would most likely reduce the broadcast operations?
MediumOther domains
All DEA-C01 exam domains
Frequently asked questions
- What does the Data Store Management domain cover on the DEA-C01 exam?
- Candidates must configure S3 lifecycle policies, DynamoDB capacity and GSIs, Glue Data Catalog partitions, and Lake Formation grants. Get S3 storage class transitions and Lake Formation column-level permissions right, since most scenario questions hinge on least-privilege access and cost-optimal storage.
- How many questions are in this domain?
- This page lists all 358 Data Store Management questions in the DEA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Store Management questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.