Amazon Web Services · Free Practice Questions · Last reviewed May 2026
24real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
34% of exam · 6 sample questions below
A data engineer needs to ingest streaming data from an IoT fleet into Amazon S3 for near-real-time analytics. The data volume is approximately 5 GB per hour, and each event is less than 1 KB. Which AWS service should be used as the ingestion endpoint?
AWS IoT Core
AWS IoT Core ingests high-volume, small-payload device telemetry and can route it directly to Amazon S3 via rules, meeting the 5 GB per hour near-real-time requirement. Its native MQTT support suits sub-1 KB events from a fleet, unlike services designed for batch or large-object transfer.
AWS DataSync
Amazon AppFlow
Amazon Kinesis Data Streams
A company uses AWS Glue ETL jobs to transform data from Amazon S3 to Amazon Redshift. The job reads JSON files, applies schema mapping, and writes to a Redshift table. Recently, the job started failing with memory errors. The data volume has increased tenfold. Which approach should a data engineer take to resolve this issue with minimal code changes?
Switch from Spark to Python Shell job type.
Implement batch processing with smaller file sizes.
Increase the number of DPUs allocated to the Glue job.
Glue allocates memory per executor, so raising the DPU count adds workers and distributes the tenfold-larger dataset across more memory, resolving the out-of-memory failures. It requires only a job configuration change, meeting the minimal-code-change constraint.
Use Redshift Spectrum to query data directly from S3.
A financial services company processes real-time stock trade data. They use Amazon Kinesis Data Streams with a shard count of 5, each shard receiving about 500 records per second. The consumer application uses the Kinesis Client Library (KCL) with DynamoDB for checkpointing. Lately, some records are being processed multiple times. What is the most likely cause?
The consumer application is crashing and restarting, causing re-processing of records.
Frequent crashes force the KCL to resume from the last DynamoDB checkpoint, replaying every record consumed after it — at-least-once delivery guarantees duplicates on restart. With five shards at 500 records per second, each restart re-processes a substantial backlog, matching the observed duplicate processing.
The Kinesis stream's iterator age is exceeding the retention period.
The DynamoDB table used for checkpointing is throttling write requests.
The record size exceeds the 1 MB API limit, causing retries.
A data engineering team needs to transform CSV files stored in Amazon S3 into Parquet format using AWS Glue. The files are partitioned by date and are updated hourly. Which AWS Glue feature should be used to automatically detect the schema and partition structure?
AWS Glue Crawler
AWS Glue Crawler scans the S3 data, infers the CSV schema, and registers partition structure in the Glue Data Catalog. Scheduled hourly, it keeps metadata current as new date partitions arrive, enabling the ETL job to read and convert files to Parquet.
AWS Glue DataBrew
AWS Lake Formation
Amazon Athena
An e-commerce company ingests clickstream data from their website into Amazon S3. The data is in JSON format, and each file is about 10 MB. They need to transform the data into a columnar format for analytics and load it into Amazon Redshift nightly. The transformation should be cost-effective and require minimal operational overhead. Which approach meets these requirements?
Use AWS Glue ETL job to convert to Parquet and load into Redshift.
AWS Glue ETL jobs convert JSON to Parquet on serverless Spark infrastructure, eliminating cluster provisioning and satisfying the minimal-operational-overhead constraint. Parquet's columnar layout and compression reduce Redshift storage and scan costs, meeting the cost-effectiveness requirement for the nightly 10 MB file transformations.
Use Amazon Redshift COPY command to load JSON directly.
Use Amazon EMR with Spark to transform and load data.
Use AWS Lambda to transform each file and write to Redshift.
A company uses AWS Database Migration Service (DMS) to continuously replicate data from an on-premises Oracle database to Amazon S3 in Parquet format. The replication is used for near-real-time analytics. Recently, the DMS task started failing with an error indicating insufficient memory. The source database is large (2 TB). What should a data engineer do to resolve this issue while minimizing changes to the existing architecture?
Change the target format to JSON to reduce memory usage.
Split the DMS task into multiple smaller tasks.
Use Change Data Capture (CDC) only, without full load.
Increase the DMS replication instance size.
DMS memory exhaustion during large-table replication is resolved by scaling the replication instance, adding RAM and CPU. This satisfies the 2 TB source constraint while preserving the existing task, endpoints and Parquet target architecture unchanged.
Want more Data Ingestion and Transformation practice?
Practice this domain22% of exam · 6 sample questions below
A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is consumed by a custom consumer application that writes to Amazon S3 every 5 minutes. The consumer is falling behind and processing lag is increasing. Which action is MOST effective to reduce the lag?
Switch to Amazon Kinesis Data Firehose to deliver data directly to S3
Increase the batch size of records written to S3
Increase the number of shards in the Kinesis stream
A Kinesis stream's throughput ceiling is set by shard count: each shard provides 1 MB/s or 1,000 records/s ingest and 2 MB/s egress. Adding shards raises parallel capacity so the consumer can drain the backlog faster, directly reducing processing lag.
Reduce the retention period of the stream
A data team runs a daily AWS Glue ETL job that processes data from an Amazon Redshift cluster and writes results to Amazon S3. The job completes successfully but takes 2 hours longer than expected. The job uses the JDBC connection to Redshift. The Redshift cluster is 4 dc2.large nodes. The Glue job has 10 workers of type G.1X. Which change would MOST likely reduce the job duration?
Use Redshift Spectrum to query data directly from S3
Use the S3 staging option in the Glue connection to unload data from Redshift to S3 first
The JDBC connector reads Redshift row-by-row through the driver, which is slow at this scale. Using the S3 staging option runs a Redshift UNLOAD to S3 in parallel, then Glue reads S3 directly, removing the JDBC bottleneck and cutting job duration.
Increase the Redshift cluster size to 8 nodes
Increase the number of Glue workers to 20
A company uses Amazon S3 to store raw data and runs AWS Glue ETL jobs to transform it into Parquet. The data is then queried using Amazon Athena. Queries are slow and expensive due to high scan volumes. Which THREE design changes can improve query performance and reduce costs? (Select THREE.)
Increase the number of files by reducing file size to 1 MB
Convert the data to a columnar format like Parquet or ORC if not already
Columnar formats such as Parquet or ORC store data by column, so Athena reads only the columns referenced in a query instead of every field. This directly cuts bytes scanned, which is the billed metric, lowering both query latency and cost for the existing Glue-produced data.
Compress the data using a splittable compression format like Snappy
Snappy is splittable, so Athena and Glue can parallelise reads across a single compressed file rather than one worker per file. Combined with reduced bytes on disk, this lowers scan volume and cost while preserving parallelism, unlike non-splittable gzip.
Use bucketing on high-cardinality columns
Partition the data by commonly filtered columns such as date or region
Partitioning by frequently filtered columns such as date or region lets Athena prune partitions via the metastore, reading only matching S3 prefixes. This directly reduces bytes scanned, the cost driver, and speeds queries that filter on those columns.
A data engineer runs a Spark job on Amazon EMR that reads data from Amazon S3 and writes results back to S3. The job fails with an 'S3AccessDenied' error. The engineer verifies that the IAM role attached to the EMR cluster has s3:GetObject and s3:PutObject permissions on the relevant buckets. What is the MOST likely cause of the error?
S3 Transfer Acceleration is not enabled on the bucket.
EMRFS consistent view is not configured.
The S3 bucket is in a different AWS Region than the EMR cluster.
The IAM role does not have s3:ListBucket permission on the bucket.
Spark's S3A filesystem lists the bucket or prefix before reading and writing objects, and that listing call requires s3:ListBucket on the bucket resource. GetObject and PutObject alone are insufficient, so the missing ListBucket permission causes the S3AccessDenied failure.
A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Redshift. The migration is successful, but after a few days, data in Redshift becomes inconsistent with the source due to ongoing changes. The company needs to keep Redshift synchronized with minimal latency. Which approach should the data engineer use?
Configure DMS with ongoing replication using change data capture (CDC).
Ongoing replication with CDC lets AWS DMS continuously apply source Oracle changes to Redshift, satisfying the requirement to keep Redshift synchronised with minimal latency after the initial full load. Without CDC, only a one-time migration occurs, so later source changes never propagate.
Use Amazon Redshift COPY with S3 staging and AWS Lambda triggers.
Schedule a full DMS load every night.
Set up Amazon Redshift Spectrum to query the Oracle database directly.
A data engineer notices that an Amazon Kinesis Data Firehose delivery stream is failing to deliver data to an Amazon S3 bucket. The CloudWatch metrics show 'DeliveryToS3.Success' is 0 and 'S3.BucketExists' is 1. What is the MOST likely cause?
The S3 bucket has an ACL that denies access to Firehose.
The Firehose delivery stream Lambda transformation function is failing.
The IAM role for Firehose lacks s3:PutObject permission.
S3.BucketExists being 1 confirms the bucket is reachable, so the failure lies in authorisation. Firehose requires s3:PutObject in its IAM role to write objects; without it, every delivery attempt is denied and DeliveryToS3.Success stays at 0.
The S3 bucket does not exist.
Want more Data Operations and Support practice?
Practice this domain18% of exam · 6 sample questions below
A company uses AWS Glue to process sensitive data stored in Amazon S3. The security team requires that all data in transit between AWS Glue and S3 be encrypted. Which configuration should be used to meet this requirement?
Use an S3 bucket policy that denies requests not using HTTPS.
An S3 bucket policy denying requests where aws:SecureTransport is false enforces TLS on every request, so Glue-to-S3 traffic cannot fall back to plain HTTP. This satisfies the in-transit encryption requirement at the bucket level, covering all clients including AWS Glue.
Use an AWS KMS key to encrypt the data before uploading to S3.
Configure AWS Glue to use SSL by setting the 'ssl' parameter to 'true'.
Enable default encryption on the S3 bucket using SSE-S3.
A company uses Amazon Redshift to store customer data. The security team requires that all queries are logged for auditing purposes. Which step should be taken to meet this requirement? (Select ONE.)
Enable AWS CloudTrail database audit logging.
Use AWS CloudTrail to log Redshift API calls.
Enable logging on the Redshift security group.
Enable VPC Flow Logs for the Redshift cluster.
Enable Amazon Redshift audit logging to an S3 bucket.
Audit logging captures connection, user, and query activity, then delivers it to Amazon S3 for durable retention. This directly satisfies the security team's requirement that all queries are logged for auditing, since S3 provides the persistent, reviewable store auditors need.
A company is using Amazon EMR to process data stored in Amazon S3. The S3 bucket is configured with a bucket policy that denies access unless the request includes a specific tag. The EMR cluster's IAM role has s3:GetObject permission. However, the EMR job fails to read data from S3. What is the most likely cause?
The bucket policy is not attached to the EMR role.
The EMR cluster is not in the same account as the S3 bucket.
The IAM role does not have a condition that matches the required tag.
The bucket policy requires a tag, and the role must have a matching condition.
The EMR role does not have s3:GetObject permission.
A company uses AWS Lake Formation to manage data lake permissions. The data lake contains sensitive customer data in the 'customer' database. The security team wants to ensure that only users with a specific tag 'access_level=analyst' can query the 'customer' table. Which combination of steps should the data engineer take to enforce this?
In Lake Formation, create an LF-tag 'access_level' with values 'analyst' and 'admin'. Grant 'SELECT' permission on the 'customer' table to the tag value 'analyst'. Associate the LF-tag with the 'customer' table.
This uses Lake Formation TBAC to restrict access based on the user's tag.
Create an IAM policy that conditionally allows 'glue:GetTable' based on the tag 'access_level=analyst'.
Apply a bucket policy on the S3 location of the 'customer' table that allows access only if the request carries the tag 'access_level=analyst'.
Use Lake Formation column-level filters to restrict access to columns based on the tag 'access_level=analyst'.
A data engineer is configuring AWS Glue jobs to access data stored in Amazon S3. The data is encrypted using server-side encryption with AWS KMS (SSE-KMS). The Glue job needs to read and write data to the S3 bucket. Which IAM policy statement should be added to the Glue job's IAM role to allow it to use the KMS key?
{"Effect":"Allow","Action":["kms:Decrypt"],"Resource":"*"}
{"Effect":"Allow","Action":["kms:Decrypt","kms:GenerateDataKey"],"Resource":"*"}
SSE-KMS requires both kms:Decrypt to read encrypted objects and kms:GenerateDataKey to obtain a data key for writing. Granting these two actions on the key resource satisfies the stem's read-and-write requirement, since Glue cannot access KMS-encrypted S3 data with either permission missing.
{"Effect":"Allow","Action":["kms:Decrypt","kms:ReEncrypt"],"Resource":"*"}
{"Effect":"Allow","Action":["kms:Decrypt","kms:Encrypt"],"Resource":"*"}
A company is building a data pipeline that ingests sensitive customer data from an on-premises database into Amazon S3 using AWS DMS. The data must be encrypted at rest in S3 and in transit. The security team requires that the encryption keys be managed by the company (not AWS). Which TWO actions should the data engineer take to meet these requirements? (Choose TWO.)
Enable encryption at rest using the default DMS encryption settings.
Configure the S3 bucket to use server-side encryption with AWS KMS (SSE-KMS) using a customer managed key.
SSE-KMS with a customer managed key encrypts objects at rest while the company retains control over key rotation and access policy, satisfying the requirement that keys not be AWS-managed. It also complements TLS for the in-transit requirement.
Configure the S3 bucket to use server-side encryption with S3 managed keys (SSE-S3).
Enable SSL/TLS encryption on the DMS source and target endpoints.
Enabling SSL/TLS on the DMS source and target endpoints encrypts replication traffic in transit, satisfying the in-transit requirement. This is separate from the at-rest requirement, which needs customer-managed keys via SSE-KMS or SSE-C on the S3 target.
Create an AWS KMS key and use it in the DMS endpoint to encrypt data in transit.
Want more Data Security and Governance practice?
Practice this domain26% of exam · 6 sample questions below
A company has an Amazon RDS for MySQL DB instance with read replicas. The primary DB instance fails. What is the correct procedure to promote a read replica to become the new primary?
Modify the read replica to be a Multi-AZ deployment and failover will occur.
RDS automatically fails over to the read replica within 5 minutes.
Manually promote the read replica to a standalone DB instance.
Promoting a read replica manually converts it into a standalone DB instance, which is the only supported route when the primary fails without Multi-AZ. The stem specifies RDS for MySQL with read replicas, so no automatic failover exists; you must trigger promotion yourself to restore write capability.
Delete the primary and the read replica will automatically become the primary.
A company uses Amazon DynamoDB for a gaming application. They need to store player session data that expires after 24 hours. Which DynamoDB feature should they use to automatically delete expired items?
Time to Live (TTL)
Time to Live (TTL) lets you define an attribute holding an expiry timestamp; DynamoDB then deletes each item automatically once that timestamp passes, with no write capacity consumed. This directly satisfies the 24-hour automatic expiry requirement for player session data, eliminating manual cleanup jobs.
DynamoDB auto scaling
DynamoDB Streams
Point-in-time recovery
An e-commerce application uses Amazon ElastiCache for Redis to cache product catalog data. The cache currently uses lazy loading. The team wants to ensure that frequently accessed product data is always fresh. Which caching strategy should they implement?
Write-through caching
Write-through caching updates the Redis cache synchronously on every database write, so cached product data never goes stale. This directly satisfies the freshness constraint that lazy loading cannot guarantee, since lazy loading only populates entries on a miss and leaves existing values stale until eviction or expiry.
Set a TTL of 5 minutes for all cached items
Use database read replicas to serve data
Lazy loading with TTL
Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket. What is the effect of this policy?
It enforces server-side encryption for all objects written to the bucket.
It allows the DataLakeRole to read and write objects, but only over HTTPS.
The policy grants DataLakeRole s3:GetObject and s3:PutObject permissions, but the condition restricts access to requests made over TLS. Plain HTTP requests are denied, so the role can read and write objects only when the connection is encrypted via HTTPS.
It allows anonymous access to the bucket for HTTPS requests.
It denies all access to the bucket except for requests from the DataLakeRole.
Refer to the exhibit. A data engineer runs the above AWS CLI command to view the table metadata in the AWS Glue Data Catalog. The data is stored as CSV in S3 with partitions by year and month. When querying the table using Amazon Athena, no data is returned. What is the most likely cause?
The partitions have not been added to the Glue Data Catalog.
Athena reads partition locations from the AWS Glue Data Catalog. If year and month partitions were never registered, the table metadata exposes no partition paths, so queries against those partitions scan nothing and return zero rows despite the CSV objects existing in S3.
The SerDe is not compatible with CSV files.
The S3 location points to a file instead of a folder.
The column data types are incorrect for the CSV data.
Which THREE storage classes in Amazon S3 are designed for infrequently accessed data with millisecond retrieval times? (Select THREE.)
S3 Glacier Flexible Retrieval
S3 One Zone-IA
S3 One Zone-IA stores data in a single Availability Zone, delivering millisecond retrieval while suiting infrequently accessed data. It satisfies the stem's latency constraint because retrieval remains immediate, unlike Glacier tiers requiring minutes to hours. Lower durability than Standard-IA is the trade-off, but the millisecond requirement is met.
S3 Glacier Deep Archive
S3 Intelligent-Tiering
S3 Intelligent-Tiering automatically moves objects between frequent and infrequent access tiers based on changing access patterns, while always delivering millisecond retrieval. This satisfies the stem's requirement for infrequently accessed data with millisecond latency, since its infrequent access tier retains the same low-latency performance as S3 Standard.
S3 Standard-IA
S3 Standard-IA suits infrequent access while preserving millisecond latency, satisfying the stem's retrieval-speed constraint. Unlike Glacier classes, which incur minutes-to-hours restore delays, Standard-IA retrieves objects immediately from the same replicated infrastructure as S3 Standard, merely charging a lower storage rate with a retrieval fee and 30-day minimum duration.
Want more Data Store Management practice?
Practice this domainThe DEA-C01 exam has 65 questions and must be completed in 130 minutes. The passing score is 720/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 4 domains: Data Ingestion and Transformation, Data Operations and Support, Data Security and Governance, Data Store Management. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Amazon Web Services DEA-C01 exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.