Courseiva

DEA-C01 · topic practice

Data Store Management practice questions

Data Store Management is 26% of DEA-C01 and covers choosing, configuring and securing AWS storage for analytics workloads. Expect scenario questions on S3 storage classes and lifecycle rules, DynamoDB capacity and indexes, Lake Formation permissions, Glue Data Catalog tables and partitions, and KMS encryption choices across Redshift, RDS and S3.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Data Store Management

What the exam tests

What to know about Data Store Management

Candidates must configure S3 lifecycle policies, DynamoDB capacity and GSIs, Glue Data Catalog partitions, and Lake Formation grants. Get S3 storage class transitions and Lake Formation column-level permissions right, since most scenario questions hinge on least-privilege access and cost-optimal storage.

Selecting S3 storage classes, lifecycle transitions and Intelligent-Tiering for cost and access patterns

Configuring DynamoDB partition keys, GSIs/LSIs, on-demand versus provisioned capacity and DynamoDB Streams

Managing Glue Data Catalog databases, tables, partitions and crawler classification of S3 data

Applying Lake Formation grants, S3 bucket policies and KMS keys for fine-grained data access control

Watch out for

Common Data Store Management exam traps

  • ▸Choosing S3 Glacier Deep Archive for data needing millisecond retrieval, ignoring the hours-long restore time and retrieval charges
  • ▸Using a low-cardinality DynamoDB partition key, causing hot partitions and throttling instead of even read/write distribution
  • ▸Granting IAM permissions on S3 but forgetting Lake Formation table grants, so Athena and Redshift Spectrum queries still fail

Practice set

Data Store Management questions

20 questions · select your answer, then reveal the explanation

A data engineering team uses Amazon Redshift for analytics. They notice that queries on a large fact table are slow. The table is distributed using DISTSTYLE ALL. Which design change would most likely improve query performance?

Which TWO actions should a data engineer take to encrypt data at rest in an Amazon S3 bucket? (Select TWO.)

A data engineer is designing a data lake on Amazon S3. The data must be immutable and support high-throughput streaming ingestion. Which THREE features should the engineer consider? (Select THREE.)

A company uses Amazon DynamoDB with on-demand capacity. They notice higher than expected costs due to a sudden spike in read traffic from a reporting job. The reporting job scans the entire table daily. What is the most cost-effective way to reduce costs while maintaining the same reporting output?

A data engineer has set up an Amazon S3 lifecycle policy to transition objects to Glacier Instant Retrieval after 30 days. After 60 days, objects should transition to Deep Archive. However, objects are not transitioning to Deep Archive. What is the most likely cause?

A data engineer attaches the above IAM policy to an IAM user. The user tries to download an object from my-bucket using the AWS CLI without specifying SSE headers. The object is stored with SSE-S3. Will the download succeed?

Exhibit

Refer to the exhibit.

```
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "s3:GetObject",
      "Resource": "arn:aws:s3:::my-bucket/*",
      "Condition": {
        "StringEquals": {
          "s3:x-amz-server-side-encryption": "AES256"
        }
      }
    }
  ]
}
```

A company uses Amazon DynamoDB for a gaming leaderboard. The table has a partition key of 'game_id' and a sort key of 'score'. The read capacity is provisioned at 1000 RCUs. During peak hours, users report high latency when querying the top 10 scores for a specific game. The DynamoDB metrics show ConsumedReadCapacityUnits averaging 800 but occasional throttling. What is the most likely cause and solution?

A data engineer runs the above AWS CLI command and receives the output. The object is part of an S3 Lifecycle policy that transitions objects to Glacier Instant Retrieval after 30 days. The object was created on January 1, 2023. Why is the object still in STANDARD_IA storage class?

Network Topology
$ aws s3api head-objectbucket my-data-lakekey logs/2023/01/01/app.logRefer to the exhibit."LastModified": "2023-01-02T00:00:00Z","ContentLength": 1048576,"ETag": "\"abc123def456\"","VersionId": "null","ContentType": "application/octet-stream","Metadata": {"x-amz-meta-original-timestamp": "2023-01-01T12:00:00Z"},"StorageClass": "STANDARD_IA","Restore": "ongoing-request="false""

A data engineer is designing a data lake on Amazon S3 for analytics. The data includes sensitive PII that must be encrypted at rest. The company requires that the encryption keys be managed by the company's own hardware security module (HSM) and rotated every 90 days. Which TWO options meet these requirements? (Choose TWO.)

A data engineering team is designing a data lake on Amazon S3 with a folder structure that separates raw, transformed, and curated data. The team needs to implement lifecycle policies to minimize storage costs while ensuring that data in the 'raw' zone is retained for 90 days before being moved to Amazon S3 Glacier Deep Archive. Additionally, data in the 'curated' zone should be deleted after 365 days. What is the MOST cost-effective way to achieve these requirements?

A data engineer is setting up an Amazon Redshift cluster for a data warehouse. The cluster will store historical sales data and support complex analytical queries. To optimize query performance and manage storage, the engineer needs to choose appropriate distribution styles and sort keys for a large fact table 'sales' and several dimension tables. Which TWO of the following design decisions are BEST practices?

A data engineer ran the above CLI command to describe an Amazon DynamoDB table named 'Orders'. The table has a key schema with 'OrderID' as the partition key and 'CustomerID' as the sort key. The table currently has no items. The engineer wants to add a new attribute 'OrderDate' and then query all orders for a specific customer within a date range. Which of the following actions is the MOST efficient approach to support this query pattern?

Network Topology
aws dynamodb describe-tabletable-name OrdersRefer to the exhibit."Table": {"AttributeDefinitions": ["AttributeName": "OrderID","AttributeType": "S"},"AttributeName": "CustomerID",],"TableName": "Orders","KeySchema": ["KeyType": "HASH""KeyType": "RANGE""TableStatus": "ACTIVE","ProvisionedThroughput": {"ReadCapacityUnits": 5,"WriteCapacityUnits": 5"TableSizeBytes": 0,"ItemCount": 0

Arrange the steps to set up cross-region replication for an S3 bucket.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4
5Step 5

A company uses Amazon S3 to store sensitive customer data. The security team requires that all objects uploaded to a specific bucket be encrypted at rest using AWS KMS with a customer managed key. Which bucket policy statement should be applied to enforce this requirement?

A company runs an Amazon RDS for MySQL database. The database experiences high write latency during peak hours. The data engineer notices that the WriteIOPS metric is consistently at the provisioned limit. Which action would most effectively reduce write latency without increasing costs?

A data engineer is configuring Amazon S3 Lifecycle policies to transition objects between storage classes. The data is accessed frequently for the first 30 days, then rarely for the next 90 days, after which it must be archived. The engineer wants to minimize costs while ensuring immediate retrieval for the first 30 days. Which lifecycle policy should the engineer implement?

A data engineer needs to store a large number of small files (each a few KB) from IoT sensors. The data is written once and never modified. The primary requirement is high write throughput and low latency for writes. Which storage solution is most suitable?

A company uses Amazon Redshift for analytics. A data engineer notices that queries are slow due to high disk usage on the compute nodes. The engineer needs to reclaim disk space without interrupting ongoing queries. Which action should the engineer take?

A data engineer is designing a data lake on Amazon S3. The data lake must support both batch and streaming ingestion. Which TWO AWS services can ingest data directly into S3? (Choose TWO.)

A company is migrating an on-premises Apache Hadoop cluster to Amazon EMR. The data is stored in HDFS and must be moved to Amazon S3. Which THREE considerations are important when designing the migration? (Choose THREE.)

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Data Store Management sessions

Start a Data Store Management only practice session

Every question in these sessions is drawn from the Data Store Management domain — nothing else.

Related practice questions

Related DEA-C01 topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the DEA-C01 exam test about Data Store Management?
Candidates must configure S3 lifecycle policies, DynamoDB capacity and GSIs, Glue Data Catalog partitions, and Lake Formation grants. Get S3 storage class transitions and Lake Formation column-level permissions right, since most scenario questions hinge on least-privilege access and cost-optimal storage.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Data Store Management questions in a focused session?
Yes — the session launcher on this page draws every question from the Data Store Management domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other DEA-C01 topics?
Use the topic links above to move to related areas, or go back to the DEA-C01 question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the DEA-C01 exam covers. They are not copied from any real exam or dump site.