Courseiva

CCNA Data Store Management Questions

75 of 358 questions · Page 4/5 · Data Store Management · Answers revealed

226
MCQhard

A company is designing a data lake on Amazon S3. The data includes personal identifiable information (PII). The data engineer must ensure that only authorized users can access the data, and that access is logged for auditing. Which combination of services should the data engineer use?

A.S3 bucket policies with IAM policies and AWS CloudTrail with data events
B.Amazon S3 access points and VPC endpoints
C.Amazon Macie to discover PII and S3 Object Lock to prevent deletion
D.AWS KMS to encrypt data and AWS CloudTrail to log access
AnswerA

S3 bucket policies and IAM policies together enforce least-privilege authorisation on the PII objects, while CloudTrail data events capture object-level API activity such as GetObject, satisfying the auditing requirement that management events alone would not record.

Why this answer

S3 bucket policies combined with IAM policies provide fine-grained access control to restrict who can access the data, while AWS CloudTrail with data events logs every S3 object-level operation (e.g., GetObject, PutObject) for auditing. This combination directly meets the requirements of authorized access and logging for PII data.

Exam trap

The trap here is that candidates often assume CloudTrail automatically logs all S3 operations, but it only logs management events by default; data events must be explicitly enabled, and many overlook this distinction when designing for auditing.

How to eliminate wrong answers

Option B is wrong because Amazon S3 access points and VPC endpoints control network-level access and simplify bucket management, but they do not provide the required logging of data access for auditing. Option C is wrong because Amazon Macie discovers and classifies PII, and S3 Object Lock prevents deletion or overwriting, but neither service enforces access control or logs data access events. Option D is wrong because AWS KMS encrypts data at rest, which protects confidentiality but does not control who can access the data, and while AWS CloudTrail logs API calls, it does not log data events by default; without enabling data event logging, object-level access (e.g., reading a file) is not recorded.

227
MCQhard

A data engineer is troubleshooting an access denied error when an AWS Lambda function tries to decrypt an object encrypted with the KMS key 'abc123'. The Lambda function's execution role has the above policy attached. What is the likely cause of the error?

A.The Deny statement blocks all decrypt requests
B.The Lambda function does not have permission to call kms:GenerateDataKey
C.The KMS key policy does not grant the Lambda role decrypt permission
D.The IAM policy does not include kms:Decrypt permission
AnswerC

The execution role's identity policy only permits the Lambda to request decryption; the KMS key policy must also allow that principal. If the key policy omits the Lambda role, KMS denies the Decrypt call, producing the access denied error despite the attached policy.

Why this answer

The error occurs because KMS requires both the IAM policy and the key policy to grant the necessary permissions. While the IAM policy attached to the Lambda execution role includes kms:Decrypt, the KMS key policy for 'abc123' does not explicitly grant the Lambda role permission to call kms:Decrypt. Since KMS key policies act as a resource-based policy, they must allow the principal (the Lambda role) to perform the action; otherwise, the request is denied even if the IAM policy allows it.

Exam trap

The trap here is that candidates assume IAM permissions alone are sufficient for KMS operations, overlooking that KMS key policies must explicitly grant access to the IAM role, which is a common source of access denied errors in cross-account or cross-service scenarios.

How to eliminate wrong answers

Option A is wrong because the Deny statement in the policy only blocks decrypt requests that do not include the encryption context 'Project=Alpha', not all decrypt requests; the error is likely due to missing key policy permissions, not a blanket Deny. Option B is wrong because the error is about decrypting an object, not generating a data key; kms:GenerateDataKey is used for encryption operations, not decryption, and the Lambda function is trying to decrypt, not encrypt. Option D is wrong because the IAM policy shown in the question includes kms:Decrypt permission (the policy lists kms:Decrypt as an allowed action), so the issue is not a missing IAM permission but rather the KMS key policy not granting the Lambda role decrypt permission.

228
MCQeasy

A company uses Amazon DynamoDB as the primary data store for a gaming application. The application stores user profiles and game state. During peak hours, the application experiences throttling on writes to the UserProfiles table. The table's read capacity is underutilized. Which solution should resolve the write throttling?

A.Increase the provisioned write capacity units for the table.
B.Enable DynamoDB Accelerator (DAX) on the table.
C.Add a global secondary index (GSI) to the table.
D.Configure auto scaling for read capacity units.
AnswerA

Increasing provisioned write capacity units directly raises the table's write throughput ceiling, eliminating the throttling caused by write requests exceeding the current WCU allocation during peak hours. Since read capacity is underutilised, only write capacity needs adjustment, making this the precise fix for the stated write bottleneck.

Why this answer

Write throttling occurs when the number of write requests exceeds the provisioned write capacity units (WCUs) for the DynamoDB table. Since the read capacity is underutilized, the correct solution is to increase the provisioned WCUs to accommodate the peak write traffic. This directly addresses the capacity deficit without affecting read operations.

Exam trap

The trap here is that candidates may confuse read and write capacity solutions, such as selecting DAX (which only helps reads) or auto scaling for reads, when the issue is specifically write throttling.

How to eliminate wrong answers

Option B is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read performance, not write throughput; it does not increase write capacity or reduce write throttling. Option C is wrong because adding a global secondary index (GSI) consumes additional write capacity from the base table and can actually increase write throttling, not resolve it. Option D is wrong because auto scaling for read capacity units does not affect write throttling; write throttling requires adjusting write capacity, not read capacity.

229
MCQmedium

A data engineer is designing a data lake on Amazon S3. The data consists of sensitive personally identifiable information (PII) that must be encrypted at rest. The company requires that encryption keys be rotated every 90 days and that access to the keys be logged. Which encryption solution meets these requirements?

A.Use server-side encryption with customer-provided keys (SSE-C).
B.Use client-side encryption with a master key stored in AWS Secrets Manager.
C.Use server-side encryption with S3 managed keys (SSE-S3).
D.Use server-side encryption with AWS KMS (SSE-KMS) and enable automatic key rotation.
AnswerD

SSE-KMS provides customer-managed keys with rotation and logging via CloudTrail.

Why this answer

SSE-KMS with automatic key rotation (Option D) meets the requirements because it encrypts data at rest in S3, allows key rotation every 90 days via AWS KMS automatic rotation, and logs all key usage in AWS CloudTrail for auditing. This provides the necessary encryption, rotation, and access logging without managing keys externally.

Exam trap

The trap here is that candidates often confuse SSE-S3's automatic annual rotation with the required 90-day rotation, or assume SSE-C or client-side encryption can meet logging and rotation requirements without realizing they lack native AWS rotation and auditing capabilities.

How to eliminate wrong answers

Option A is wrong because SSE-C requires the customer to provide and manage their own encryption keys, and AWS does not support automatic key rotation or logging of key access for customer-provided keys. Option B is wrong because client-side encryption encrypts data before it reaches S3, but storing the master key in AWS Secrets Manager does not provide automatic key rotation every 90 days (Secrets Manager rotation is configurable but not native to KMS key rotation) and does not log key usage in CloudTrail as KMS does. Option C is wrong because SSE-S3 uses S3-managed keys that are rotated annually (not every 90 days) and do not provide granular access logging for key usage.

230
MCQmedium

A data engineer is building a data lake on Amazon S3 and must choose the optimal file format for a dataset that is queried by Amazon Athena. The queries typically select a few columns from wide tables containing hundreds of columns, and the data volume is in terabytes. The engineer wants to minimize query scan costs and improve performance. Which file format should the engineer use?

A.Apache Avro
B.CSV
C.Apache Parquet
D.JSON
AnswerC

Parquet is a columnar format that stores data by column, enabling Athena to read only the columns referenced in a query. This drastically reduces the amount of data scanned, lowering costs and improving performance for wide tables where only a few columns are accessed. It also supports efficient compression and encoding, further reducing storage and scan volume, making it ideal for this analytical workload.

Why this answer

Apache Parquet is a columnar storage format that allows Amazon Athena to read only the columns needed for a query, significantly reducing the amount of data scanned. For wide tables where queries access a small subset of columns, this column pruning leads to lower costs and faster performance. Parquet also offers efficient compression, further reducing storage and scan volume, making it the best choice for this scenario.

Exam trap

The trap here is assuming that any columnar format is equally optimal, but Parquet specifically excels for Athena due to its widespread support and predicate pushdown capabilities.

231
MCQhard

A CloudFormation template includes this IAM policy for a cross-account S3 upload use case. What is the purpose of the condition?

A.To enforce server-side encryption with KMS.
B.To limit the size of objects that can be uploaded.
C.To restrict uploads to only a specific AWS account.
D.To ensure that uploaded objects grant full control to the bucket owner.
AnswerD

The condition checks the canned ACL on uploaded objects, ensuring the bucket owner receives full control. This satisfies the cross-account upload scenario, where objects written by another account would otherwise remain owned solely by the uploading account.

Why this answer

The condition in the IAM policy uses the `s3:x-amz-acl` key with a value of `bucket-owner-full-control`. This ensures that any object uploaded to the S3 bucket explicitly grants the bucket owner full control over the object, overriding the default behavior where the uploading account retains ownership. This is critical in cross-account uploads to prevent the uploading account from retaining exclusive access to the objects.

Exam trap

The trap here is that candidates confuse the `s3:x-amz-acl` condition key with account-level restrictions or encryption settings, when in fact it specifically controls the Access Control List (ACL) applied to the uploaded object.

How to eliminate wrong answers

Option A is wrong because server-side encryption with KMS is enforced using the `s3:x-amz-server-side-encryption-aws:kms` condition key, not the `s3:x-amz-acl` key. Option B is wrong because object size limits are enforced using the `s3:content-length-range` condition key, not the ACL-related condition shown. Option C is wrong because restricting uploads to a specific AWS account is done using the `aws:SourceAccount` or `aws:SourceArn` condition keys, not the `s3:x-amz-acl` key which controls object ACL permissions.

232
MCQmedium

A data engineer runs the above command and gets the output. What does the 'MFADelete' setting imply?

A.Any modification to an object requires MFA.
B.To permanently delete a version of an object, the user must provide MFA.
C.MFA is required for all read operations as well.
D.All operations on the bucket require MFA authentication.
AnswerB

With MFADelete enabled on a versioned bucket, deleting a specific object version requires an MFA token in the request; without it, the delete is denied. This protects against accidental or malicious permanent version removal, satisfying the stem's requirement to understand what the setting enforces.

Why this answer

The 'MFADelete' setting on an S3 bucket versioning configuration requires multi-factor authentication to permanently delete an object version. This means that when a user issues a DELETE request with a version ID (a permanent delete), they must include a valid MFA token in the request headers. It does not apply to creating new versions, reading objects, or other operations.

Exam trap

The DEA-C01 exam often tests the distinction between 'MFA Delete' (which only applies to permanent version deletion) and general MFA enforcement on all bucket operations, leading candidates to overgeneralize the scope of the setting.

How to eliminate wrong answers

Option A is wrong because 'MFADelete' does not require MFA for any modification (e.g., PUT, POST, or COPY operations); it only applies to permanent deletes of specific versions. Option C is wrong because read operations (GET, HEAD) are never subject to MFA requirements under this setting. Option D is wrong because 'MFADelete' is a versioning-specific sub-setting and does not enforce MFA on all bucket operations, only on permanent version deletion.

233
MCQhard

A company runs an Amazon RDS for PostgreSQL instance that stores financial data. The company requires point-in-time recovery (PITR) with a retention period of 35 days. Additionally, the company needs to create a new database from a specific snapshot every night for testing. Which combination of actions should the data engineer take to meet these requirements?

A.Enable automated backups with a 35-day retention period and create a manual snapshot each night for testing.
B.Create a read replica and promote it to a new instance for testing each night.
C.Enable Multi-AZ and use the standby instance for testing.
D.Disable automated backups to reduce storage costs and take manual snapshots with 35-day retention.
AnswerA

Automated backups provide continuous PITR up to 35 days, meeting the retention requirement. Manual snapshots are independent of the backup retention window and can be created nightly, giving a stable source for the test database without affecting PITR.

Why this answer

Automated backups in Amazon RDS for PostgreSQL support a maximum retention period of 35 days, which satisfies the PITR requirement. Additionally, creating a manual snapshot each night provides a stable, independent copy for testing without interfering with the automated backup schedule or the source database's performance.

Exam trap

The trap here is that candidates often confuse the purpose of Multi-AZ standby instances (which are not directly usable for testing) or assume that manual snapshots alone can provide PITR, but automated backups are strictly required for point-in-time recovery in RDS.

How to eliminate wrong answers

Option B is wrong because a read replica is designed for read scaling and high availability, not for creating a nightly test database; promoting a read replica each night would disrupt replication and require re-creating the replica, which is inefficient and does not meet the PITR retention requirement. Option C is wrong because Multi-AZ provides high availability and automatic failover, but the standby instance is not directly accessible for testing; it cannot be used to create a new database without promoting it, which would break the Multi-AZ configuration. Option D is wrong because disabling automated backups eliminates the ability to perform point-in-time recovery (PITR), and manual snapshots alone do not support PITR; automated backups are required for transaction log retention and restore to any point within the retention window.

234
MCQeasy

A data engineer needs to store log files from multiple applications in a centralized location. The logs are generated in JSON format and each log entry is about 1 KB. The engineer needs to query the logs occasionally using SQL-like queries. Which AWS service is most appropriate?

A.Amazon DynamoDB
B.Amazon Redshift
C.Amazon Athena with data stored in S3
D.Amazon RDS for MySQL
AnswerC

Athena queries JSON in S3 directly using standard SQL, charging only per query scanned, which suits occasional log analysis. S3 provides the centralised, durable store for the 1 KB JSON entries, so no database loading or cluster is required.

Why this answer

Amazon Athena is the most appropriate service because it allows you to query log files stored in S3 directly using standard SQL, without needing to load or transform the data. Since the logs are in JSON format and each entry is about 1 KB, Athena's schema-on-read approach works perfectly for occasional SQL-like queries, and you only pay for the data scanned per query, making it cost-effective for infrequent access.

Exam trap

The trap here is that candidates often choose Amazon Redshift or RDS because they think 'SQL-like queries' require a traditional database, overlooking Athena's ability to query data directly in S3 without loading it, which is a key serverless pattern for log analytics.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for low-latency, high-throughput access patterns, not for ad-hoc SQL-like queries on large volumes of log data, and it would require schema design and provisioning. Option B is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for complex analytical queries on structured, transformed data, which is overkill and costly for occasional log queries on JSON files stored in S3. Option D is wrong because Amazon RDS for MySQL is a relational database that requires schema definition, data loading, and ongoing management, making it unsuitable for storing raw JSON log files directly without ETL, and it lacks the serverless, pay-per-query model for infrequent access.

235
MCQeasy

A company is using Amazon EMR to process large datasets stored in Amazon S3. The data engineer wants to reduce the time it takes to read data from S3 by optimizing the data format. Which file format should the engineer recommend?

A.CSV
B.Parquet
C.ORC
D.JSON
AnswerB

Parquet is columnar and compressed, so EMR reads only the columns referenced by the query rather than scanning entire rows. This column pruning plus predicate pushdown and efficient encoding cuts the volume of data transferred from S3, directly reducing read time.

Why this answer

Parquet is the correct choice because it is a columnar storage format that significantly reduces the amount of data read from Amazon S3 during analytical queries. By storing data column-wise, Parquet enables predicate pushdown and compression, which minimizes I/O and speeds up data processing in Amazon EMR, especially for large datasets.

Exam trap

The trap here is that candidates often assume ORC is the default or preferred format for all big data engines. However, for Amazon EMR, Parquet is generally recommended because of its superior performance with Spark and its ability to handle complex nested data structures efficiently.

How to eliminate wrong answers

Option A is wrong because CSV is a row-oriented text format that requires full file scans and offers no compression or predicate pushdown, leading to slower reads. Option C is wrong because ORC is also a columnar format optimized for Hive workloads, but it is not natively as performant with Spark and EMR as Parquet, and the question asks for the best recommendation for EMR. Option D is wrong because JSON is a row-oriented, self-describing format that is verbose and lacks efficient compression or columnar access patterns, resulting in high I/O and slower processing.

236
Multi-Selecthard

Which TWO are benefits of using Amazon S3 Object Lock? (Choose TWO.)

Select 2 answers
A.Helps meet regulatory requirements for write-once-read-many (WORM) storage.
B.Encrypts objects at rest using AWS KMS.
C.Prevents objects from being deleted or overwritten for a fixed time.
D.Automatically transitions objects to lower-cost storage classes.
E.Enables automatic versioning of objects.
AnswersA, C

Object Lock provides WORM semantics, meaning objects written once cannot be altered or removed during retention. This directly satisfies the stem's regulatory compliance benefit, since auditors accept S3 Object Lock as evidence of tamper-proof, immutable record retention.

Why this answer

Amazon S3 Object Lock helps meet regulatory requirements for write-once-read-many (WORM) storage by allowing you to set retention periods and legal holds on objects. This ensures that data cannot be deleted or overwritten for a specified duration, which is a common requirement for compliance frameworks such as SEC Rule 17a-4 or FINRA.

Exam trap

The trap here is that candidates confuse S3 Object Lock with S3 Versioning or S3 Lifecycle policies, mistakenly thinking Object Lock handles encryption or storage tier transitions, when in reality it is solely focused on preventing object deletion or overwrite for compliance-driven WORM scenarios.

237
MCQhard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak hours. The cluster uses a single node type and has no concurrency scaling enabled. Analysis shows that many long-running queries are queued behind short ad-hoc queries, causing delays for critical reports. The engineer needs to ensure that critical reports run promptly without affecting ad-hoc queries. Which solution meets these requirements?

A.Create a workload management (WLM) queue with a higher priority for critical reports and assign them to that queue.
B.Use short query acceleration (SQA) to run short queries in a dedicated space.
C.Increase the cluster size by adding more nodes to reduce overall query time.
D.Enable concurrency scaling to automatically add transient clusters for queued queries.
AnswerA

Amazon Redshift WLM allows you to create separate queues with different priorities. By assigning critical reports to a high-priority queue, they are scheduled before queries in lower-priority queues. This ensures that critical reports run promptly even when ad-hoc queries are present. This is the most direct way to prioritize specific workloads without affecting others.

Why this answer

Amazon Redshift workload management (WLM) enables you to define queues with different priorities. By routing critical reports to a high-priority queue, they are scheduled ahead of ad-hoc queries, ensuring timely execution. This approach directly addresses the need to prioritize specific workloads without impacting others.

Exam trap

The trap here is confusing concurrency scaling or SQA with workload prioritization; these features improve throughput or short query performance but do not enforce priority for specific queries.

238
MCQhard

A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a table. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition and ensure that the load is efficient and cost-effective. Which method should the engineer use?

A.Use AWS Glue to read the Parquet data, filter by the latest date, and write to Redshift.
B.Use the COPY command with the FROM 's3://bucket/prefix/date=2023-10-01/' option to load only that partition.
C.Use Amazon Redshift Spectrum to create an external table and query the latest partition.
D.Use the COPY command with the FROM 's3://bucket/prefix' option and specify the partition column in the WHERE clause.
AnswerB

The COPY command can load data from a specific S3 prefix. By specifying the exact prefix for the desired partition, the engineer loads only that partition's files. This is efficient and cost-effective because it avoids scanning and loading unnecessary data. Parquet format is natively supported, and the load will be fast and compressed.

Why this answer

The COPY command can load data directly from a specific S3 prefix, so specifying the exact partition prefix loads only that partition efficiently. This avoids unnecessary data transfer and cost. Using a WHERE clause with COPY is invalid, and alternatives like AWS Glue or Redshift Spectrum add complexity or do not load data into Redshift tables as required.

Exam trap

The trap here is assuming that the COPY command supports filtering with a WHERE clause, when in fact you must specify the exact S3 prefix to load a subset of data.

239
MCQhard

A data engineer needs to set up a new Amazon RDS for PostgreSQL database for a production workload. The database must be highly available and resilient to a single Availability Zone failure. Which configuration should the engineer choose?

A.Single-AZ with automated backups
B.Multi-AZ deployment with one standby in a different AZ
C.Multi-AZ with two readable standbys
D.Single-AZ with a read replica
AnswerB

Multi-AZ maintains a synchronous standby replica in a separate Availability Zone, so RDS automatically fails over to it if the primary AZ fails. This directly satisfies the resilience-to-single-AZ-failure constraint, unlike a single-AZ instance or read replicas, which use asynchronous replication and do not provide automatic failover.

Why this answer

A Multi-AZ deployment for Amazon RDS PostgreSQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. This configuration provides automatic failover in the event of an AZ failure, ensuring high availability and resilience without manual intervention. The synchronous replication ensures zero data loss during failover, which is critical for production workloads.

Exam trap

The trap here is that candidates often confuse Multi-AZ with read replicas, assuming that a read replica can serve as a failover target, but in RDS PostgreSQL, read replicas are asynchronous and require manual promotion, making them unsuitable for automatic high availability against AZ failures.

How to eliminate wrong answers

Option A is wrong because a Single-AZ deployment with automated backups only protects against data loss via point-in-time recovery, but does not provide automatic failover or resilience to an AZ failure; the database becomes unavailable if the AZ goes down. Option C is wrong because Amazon RDS for PostgreSQL does not support Multi-AZ with two readable standbys; that feature is specific to Amazon RDS for Oracle and SQL Server Enterprise Edition, and PostgreSQL Multi-AZ only provides a single standby that is not readable. Option D is wrong because a Single-AZ with a read replica provides read scaling and some disaster recovery capability, but the read replica is asynchronous and does not provide automatic failover; a manual promotion is required, and the primary remains vulnerable to AZ failure.

240
MCQhard

A company is using Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key. The data engineer must ensure that only a specific IAM role can decrypt the data. Which policy should the data engineer attach to the KMS key?

A.A KMS key policy that allows the IAM role to perform kms:Decrypt
B.An IAM user policy that allows kms:Decrypt for the specific key
C.An IAM policy attached to the role that allows kms:Decrypt
D.An S3 bucket policy that denies access unless encryption is used
AnswerA

KMS key policies grant permissions to use the key.

Why this answer

KMS key policies are the primary mechanism for controlling access to a customer-managed KMS key. By specifying the IAM role as a principal in the key policy and granting kms:Decrypt, you ensure that only that role can decrypt data encrypted with this key, regardless of any IAM policies that might otherwise allow broader access.

Exam trap

The DEA-C01 exam often tests the misconception that IAM policies alone can control KMS key access, but the correct approach is to use a KMS key policy that explicitly grants the required action to the specific principal.

How to eliminate wrong answers

Option B is wrong because an IAM user policy alone is insufficient; KMS key access requires either a key policy that explicitly grants permissions to the user/role or a grant, and IAM policies only take effect if the key policy allows IAM policy-based access (via a root principal). Option C is wrong because while an IAM policy attached to the role can allow kms:Decrypt, it will only work if the KMS key policy also permits IAM policy-based access (e.g., by allowing the root account), which is not guaranteed and does not restrict decryption to that specific role as tightly as a key policy. Option D is wrong because an S3 bucket policy that denies access unless encryption is used does not control who can decrypt data; it only enforces encryption in transit or at rest, and does not restrict decryption permissions to a specific IAM role.

241
MCQeasy

A data engineer is configuring an Amazon S3 bucket that will receive raw clickstream files from a mobile application. The engineer must ensure that the objects are protected against accidental overwrites and deletions for a defined retention period, and that the protection cannot be removed or shortened by any user, including the account root user. Which S3 feature should the engineer use?

A.S3 Object Lock in compliance mode with a retention period matching the required window.
B.S3 Versioning with a lifecycle rule that transitions objects to S3 Standard-IA after 30 days.
C.A bucket policy that denies s3:DeleteObject and s3:PutObject to all principals except the ingestion role.
D.S3 default encryption with AWS KMS customer managed keys and a restrictive key policy.
AnswerA

S3 Object Lock in compliance mode prevents any user, including the account root user, from overwriting or deleting a protected object version until the retention date passes, and the retention cannot be shortened. That exactly matches the requirement for enforceable, tamper-proof retention on the incoming clickstream objects.

Why this answer

S3 Object Lock provides write-once-read-many protection by binding a retention period to object versions. In compliance mode, no principal, including the account root user, can shorten or remove that retention, which is the only option that satisfies an absolute immutability requirement. Governance mode, by contrast, allows privileged users to bypass retention.

Exam trap

The trap here is assuming that a deny-based bucket policy or versioning provides the same guarantee as Object Lock, when both can be changed or bypassed by privileged identities.

242
MCQmedium

Refer to the exhibit. A data engineer runs the above AWS CLI command to view the table metadata in the AWS Glue Data Catalog. The data is stored as CSV in S3 with partitions by year and month. When querying the table using Amazon Athena, no data is returned. What is the most likely cause?

A.The partitions have not been added to the Glue Data Catalog.
B.The SerDe is not compatible with CSV files.
C.The S3 location points to a file instead of a folder.
D.The column data types are incorrect for the CSV data.
AnswerA

Athena reads partition locations from the AWS Glue Data Catalog. If year and month partitions were never registered, the table metadata exposes no partition paths, so queries against those partitions scan nothing and return zero rows despite the CSV objects existing in S3.

Why this answer

The AWS CLI command shown only retrieves table metadata, not partition metadata. In AWS Glue, partitions must be explicitly added to the Data Catalog via `MSCK REPAIR TABLE`, `ALTER TABLE ADD PARTITION`, or a Glue crawler. Without partition metadata, Athena cannot locate the data files under the partitioned S3 paths (e.g., `s3://bucket/year=2024/month=01/`), resulting in zero rows returned even though the table schema is defined.

Exam trap

The trap here is that candidates assume the `PARTITIONED BY` clause in the table definition automatically registers the partitions in the Glue Data Catalog, but it only defines the schema; partition metadata must be added separately.

How to eliminate wrong answers

Option B is wrong because the default SerDe for CSV in Athena (`LazySimpleSerDe`) is fully compatible with standard CSV files; no SerDe mismatch would cause zero rows. Option C is wrong because the `LOCATION` in the Glue table points to a folder (the base path), not a file; Athena expects a folder and would fail with an error if a file were specified, not silently return no data. Option D is wrong because incorrect column data types would cause query failures or data conversion errors, not an empty result set; Athena would still attempt to read the data and return rows with nulls or errors.

243
MCQmedium

Refer to the exhibit. A data engineer notices that the Redshift cluster 'mycluster' does not have automated backups beyond 7 days. However, the compliance team requires a minimum of 35 days of backup retention. What should the engineer do?

A.Change the node type to ra3.xlplus to enable automatic backups for 35 days.
B.Enable audit logging to capture changes for recovery.
C.Take manual snapshots every day and retain them for 35 days.
D.Modify the cluster's automated snapshot retention period to 35 days.
AnswerD

Redshift's automated snapshot retention is a cluster-level setting adjustable from 1 to 35 days. Modifying it to 35 days satisfies the compliance requirement directly, without manual snapshots or a new cluster, since the current 7-day default falls short.

Why this answer

Amazon Redshift allows you to modify the automated snapshot retention period for a cluster up to 35 days. The engineer can use the AWS Management Console, CLI, or API to change the `automated_snapshot_retention_period` parameter from the current 7 days to 35 days, meeting the compliance requirement without additional manual intervention.

Exam trap

The trap here is that candidates may confuse backup retention with node type capabilities or audit logging, assuming that hardware or logging features inherently extend backup duration, when in fact the retention period is a simple configuration parameter.

How to eliminate wrong answers

Option A is wrong because changing the node type to ra3.xlplus does not affect the automated backup retention period; retention is configured independently of node type. Option B is wrong because audit logging captures user activity and SQL queries for security and compliance, not for point-in-time recovery of data; it does not replace backup retention. Option C is wrong because while manual snapshots can be retained for 35 days, this approach requires daily manual effort and does not leverage the automated backup feature that is already available; modifying the automated retention period is simpler and more reliable.

244
MCQeasy

A data engineer needs to store streaming data from IoT devices for real-time analytics. The data has a fixed schema and requires low-latency queries. Which AWS service should be used?

A.Amazon DynamoDB
B.Amazon Redshift
C.Amazon S3
D.Amazon Timestream
AnswerD

Amazon Timestream is purpose-built for time-series IoT data, offering serverless auto-scaling and its memory store for low-latency queries on recent data, with magnetic store for historical retention. It satisfies the fixed-schema, real-time analytics constraint directly, unlike general-purpose databases that require manual provisioning and tuning for streaming ingest.

Why this answer

Amazon Timestream is a time-series database purpose-built for IoT and operational applications that generate large volumes of time-stamped data. It automatically manages data retention and storage tiers (memory and magnetic) to provide fast query performance for recent data and cost-effective storage for historical data, making it ideal for real-time analytics on streaming IoT data with a fixed schema.

Exam trap

AWS often tests the misconception that any database can handle time-series data equally well, but the trap here is that candidates choose DynamoDB for its low-latency reads, overlooking that Timestream is the only AWS service purpose-built for time-series workloads with native support for time-based partitioning, retention policies, and analytical functions.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for high-throughput, low-latency read/write operations on individual items, but it lacks native time-series optimizations such as automatic downsampling, interpolation, and time-based partitioning, making it less efficient for time-series queries like aggregations over time windows. Option B is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for complex analytical queries on structured and semi-structured data using SQL, but it is not optimized for real-time streaming ingestion or low-latency queries on high-frequency time-series data; its batch-oriented architecture introduces higher latency for streaming use cases. Option C is wrong because Amazon S3 is an object storage service that provides durable, scalable storage for any type of data, but it does not support real-time querying directly; querying S3 requires services like Athena or S3 Select, which add latency and are not designed for sub-second, low-latency queries on streaming data.

245
MCQhard

A data engineer is designing a data lake on Amazon S3. The data is frequently accessed by multiple analytics services, and the company needs to enforce fine-grained access control based on data tags. Which combination of AWS services should be used?

A.S3 Block Public Access settings
B.AWS Lake Formation with tag-based access control
C.S3 Access Points with bucket policies
D.S3 Object Lambda with IAM policies
AnswerB

Lake Formation tag-based access control enforces column, row and table permissions using LF-Tags across analytics services, satisfying the fine-grained, tag-driven requirement. Plain S3 bucket policies or IAM alone cannot express tag-based governance at that granularity.

Why this answer

AWS Lake Formation with tag-based access control (TBAC) is the correct choice because it provides fine-grained, attribute-based access control (ABAC) at the column, row, and cell level across a data lake on S3. By assigning LF-tags to Data Catalog resources and defining permissions based on those tags, you can enforce granular access policies that scale without managing individual user-to-resource mappings. This directly meets the requirement for tag-driven, fine-grained access for multiple analytics services.

Exam trap

The trap here is that candidates often confuse S3 Access Points (which provide network-level or prefix-level restrictions) with the fine-grained, tag-driven access control that Lake Formation TBAC uniquely offers, leading them to pick Option C despite its inability to enforce column- or row-level security based on tags.

How to eliminate wrong answers

Option A is wrong because S3 Block Public Access settings only prevent public exposure of S3 objects and do not provide any fine-grained, tag-based access control for internal users or services. Option C is wrong because S3 Access Points with bucket policies can restrict access based on VPC or IP, but they do not natively support tag-based access control at the column or row level; they operate at the bucket or prefix level only. Option D is wrong because S3 Object Lambda transforms data on read but does not enforce access control based on data tags; IAM policies attached to it cannot dynamically filter data by tags without custom code, and it lacks the centralized governance Lake Formation provides.

246
MCQeasy

A data engineer needs to store JSON documents that are frequently accessed by a low-latency web application. The data does not require complex queries, and the access pattern is primarily by a key. Which AWS service is most appropriate?

A.Amazon ElastiCache for Redis
B.Amazon S3
C.Amazon RDS for MySQL
D.Amazon DynamoDB
AnswerD

Amazon DynamoDB stores JSON as native items and retrieves them by primary key in single-digit milliseconds, satisfying the low-latency, key-based access pattern. Its schemaless design handles document data without complex query requirements, unlike relational engines or object storage, which add overhead unnecessary for this workload.

Why this answer

Amazon DynamoDB is the most appropriate service because it is a fully managed NoSQL key-value and document database designed for single-digit millisecond latency at any scale. It natively supports JSON documents and provides fast, consistent access by primary key without requiring complex query capabilities, making it ideal for low-latency web applications.

Exam trap

The trap here is that candidates often confuse ElastiCache for Redis as a persistent data store for JSON documents, but it is primarily an in-memory cache with optional persistence, not a durable, low-latency database designed for primary key access patterns.

How to eliminate wrong answers

Option A is wrong because Amazon ElastiCache for Redis is an in-memory data store primarily used for caching, session management, and real-time analytics, not for persistent storage of JSON documents that need durability and key-based access with low latency. Option B is wrong because Amazon S3 is an object storage service with higher latency (typically tens to hundreds of milliseconds) and is not optimized for frequent, low-latency key-based lookups required by a web application. Option C is wrong because Amazon RDS for MySQL is a relational database that requires predefined schemas and supports complex queries via SQL, which is overkill and adds unnecessary overhead for simple key-based access to JSON documents.

247
MCQeasy

A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a folder structure. The engineer wants to use AWS Glue crawlers to automatically infer schemas and create tables in the AWS Glue Data Catalog. The crawler should run daily to detect new files and schema changes. Which configuration should the engineer use for the crawler?

A.Set the crawler's data source to the specific S3 folder and disable the crawler schedule, running it manually when needed.
B.Set the crawler's data source to the specific S3 folder containing the CSV files, configure it to create a single schema for each S3 path, and set a daily schedule.
C.Set the crawler's data source to the specific S3 folder and configure it to create a separate table for each file, with a daily schedule.
D.Set the crawler's data source to the S3 bucket root and enable 'Update the table definition in the Data Catalog' with a schedule of daily.
AnswerB

Pointing the crawler to the specific folder limits its scope to relevant data. Configuring it to create a single schema for each S3 path groups files with the same schema into one table, which is efficient for partitioned data. A daily schedule ensures new files and schema changes are detected automatically.

Why this answer

The crawler should be pointed to the specific S3 folder to avoid scanning irrelevant data. Configuring it to create a single schema for each S3 path groups files with the same schema into one table, which is ideal for a data lake with partitioned folders. A daily schedule ensures the catalog stays up to date with new files and schema changes.

Exam trap

The trap here is choosing to create a separate table for each file, which seems granular but leads to catalog sprawl and is not how crawlers are typically used for partitioned data.

248
MCQhard

A company uses Amazon DynamoDB with on-demand capacity for a gaming application that experiences unpredictable traffic spikes. The application reads the same set of 'hot' items frequently. Users report high latency during peak hours. Which action would MOST effectively reduce read latency for the hot items?

A.Enable DynamoDB Accelerator (DAX) for the table.
B.Switch to provisioned capacity with auto-scaling.
C.Increase the read capacity units for the table.
D.Enable DynamoDB Global Tables for multi-region replication.
AnswerA

DAX provides an in-memory write-through cache for eventually consistent reads, absorbing repeated access to the same hot items and cutting microsecond-level latency. This directly addresses the unpredictable spikes and repeated hot-item reads that overwhelm on-demand capacity.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache that sits between the application and DynamoDB, providing microsecond read latency for frequently accessed items. Since the application reads the same set of 'hot' items repeatedly, DAX can serve these reads from its cache, bypassing the storage layer and reducing latency during traffic spikes without requiring any table schema changes.

Exam trap

The trap here is that candidates often confuse throughput capacity (RCUs/WCUs) with latency, assuming that increasing capacity will speed up individual reads, when in fact capacity only controls the rate of requests, not the response time per request.

How to eliminate wrong answers

Option B is wrong because switching to provisioned capacity with auto-scaling does not reduce read latency; it only manages throughput capacity based on load, but the underlying read latency from DynamoDB remains the same. Option C is wrong because increasing read capacity units (RCUs) is only applicable to provisioned capacity mode, not on-demand capacity, and even if it were, it would not reduce latency for hot items—it only increases the maximum throughput. Option D is wrong because DynamoDB Global Tables replicate data across regions for disaster recovery and low-latency reads from distant regions, but it does not reduce latency for reads within the same region; it adds complexity and cost without addressing the hot-item caching issue.

249
MCQmedium

The exhibit shows an S3 bucket policy. What is the effect of this policy?

A.Allows all S3 actions over HTTPS only.
B.Allows all S3 actions to the bucket over any protocol.
C.Denies all S3 actions to the bucket.
D.Allows only GetObject and PutObject over HTTPS.
AnswerD

Explicit allow for those actions over HTTPS; deny for HTTP.

Why this answer

The S3 bucket policy in the exhibit uses a condition key `aws:SecureTransport` set to `true`, which restricts access to HTTPS only. The `Effect` is `Allow` for `s3:GetObject` and `s3:PutObject` actions, meaning only these two actions are permitted over HTTPS. This matches option D.

Exam trap

The trap here is that candidates see the `Deny` statement and assume the entire policy denies all actions, overlooking the `Allow` statement that permits specific actions over HTTPS.

How to eliminate wrong answers

Option A is wrong because the policy does not allow all S3 actions; it explicitly allows only `s3:GetObject` and `s3:PutObject`. Option B is wrong because the policy denies all actions over non-HTTPS protocols via the `Deny` statement with `aws:SecureTransport=false`, and the `Allow` statement only permits HTTPS. Option C is wrong because the policy does not deny all S3 actions; it allows `GetObject` and `PutObject` over HTTPS, while only denying actions that do not use HTTPS.

250
MCQeasy

A data engineer is building a data lake on Amazon S3. The engineer needs to catalog metadata for data stored in Parquet format and make it queryable by Amazon Athena. The data is partitioned by year, month, and day in the S3 path. Which AWS service should the engineer use to create and manage the table definitions and partitions?

A.Amazon Athena Data Catalog
B.AWS Glue Data Catalog
C.AWS Lake Formation
D.Amazon S3 Inventory
AnswerB

AWS Glue Data Catalog is a centralized metadata repository that stores table definitions, schemas, and partition information. It integrates natively with Amazon Athena, allowing the engineer to define tables and partitions once and query them directly from Athena without additional configuration.

Why this answer

AWS Glue Data Catalog is the metadata store for Amazon Athena. It allows the data engineer to define table schemas and partitions that Athena can query directly. Using the Glue Data Catalog ensures seamless integration and simplifies data discovery and querying.

Exam trap

The trap here is thinking that Athena has its own separate data catalog, when in fact it relies exclusively on the AWS Glue Data Catalog for table metadata.

251
MCQeasy

A company stores time-series sensor data in Amazon S3. They need to query the data using SQL with minimal latency and no infrastructure management. Which service should they use?

A.Amazon Kinesis Data Analytics
B.Amazon Athena
C.Amazon Redshift
D.Amazon DynamoDB
AnswerB

Amazon Athena queries S3 data directly using standard SQL, with no servers to provision or manage. It satisfies both stated constraints: minimal latency through parallel query execution, and zero infrastructure management. Unlike Amazon Redshift, which requires cluster provisioning, Athena is serverless and pay-per-query, making it ideal for ad hoc time-series analysis.

Why this answer

Amazon Athena is the correct choice because it is a serverless interactive query service that allows you to analyze data directly in Amazon S3 using standard SQL without any infrastructure to manage. It is optimized for querying structured, semi-structured, and unstructured data stored in S3, making it ideal for time-series sensor data with minimal latency requirements.

Exam trap

The trap here is that candidates often confuse Amazon Athena with Amazon Redshift Spectrum, but the question explicitly requires 'no infrastructure management,' which eliminates Redshift; also, Kinesis Data Analytics is mistakenly chosen by those who think it can query static S3 data, but it is strictly for real-time streams.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Analytics is designed for real-time stream processing using SQL or Apache Flink, not for querying static data already stored in S3; it requires a streaming data source and incurs ongoing processing costs. Option C is wrong because Amazon Redshift is a fully managed data warehouse that requires provisioning and managing clusters, which contradicts the 'no infrastructure management' requirement; it is also overkill for simple SQL queries on S3 data and incurs higher costs for idle compute. Option D is wrong because Amazon DynamoDB is a NoSQL key-value and document database, not designed for SQL queries on S3 data; it requires data to be loaded into tables and does not support direct querying of S3 objects.

252
MCQmedium

A data engineer is designing an Amazon S3 data lake and needs to enforce schema-on-read for a dataset that is queried by Amazon Athena. The data is stored as Parquet files partitioned by year, month, and day. The engineer wants to minimize the amount of data scanned by queries that filter on a specific date range. Which approach should the engineer take?

A.Create an AWS Glue Data Catalog table with partition keys year, month, and day, and run MSCK REPAIR TABLE or use partition projection before querying in Athena.
B.Store all Parquet files in a single prefix and rely on Parquet column pruning to skip dates.
C.Convert the dataset to CSV and add a WHERE clause on the date column in every query.
D.Create an Amazon Redshift Spectrum external table and query it from Redshift instead of Athena.
AnswerA

Registering the partitions in the Data Catalog lets Athena prune partitions that do not match the date filter, so only the relevant S3 prefixes are read. Partition projection can compute partition locations from the table properties without running a repair step, which is more scalable. Either way, partition pruning is what minimizes scanned bytes.

Why this answer

Athena uses the AWS Glue Data Catalog to resolve table schemas and partition locations. When the year, month, and day columns are declared as partition keys and the partitions are either registered or projected, date filters cause Athena to read only the matching S3 prefixes. Storing everything in one prefix, converting to CSV, or moving to Redshift Spectrum does not achieve the same partition-pruning benefit.

Exam trap

The trap here is confusing Parquet column pruning and predicate pushdown with partition pruning, assuming file format alone eliminates scanning of non-matching dates.

253
MCQeasy

A data engineer needs to store semi-structured JSON transaction logs for analytics. The logs are written once and rarely accessed. The storage must be cost-effective. Which AWS service should be used?

A.Amazon S3
B.Amazon DynamoDB
C.Amazon RDS
D.Amazon Redshift
AnswerA

Amazon S3 provides durable object storage with tiered classes such as S3 Standard-IA and Glacier, matching the write-once, rarely accessed, cost-sensitive requirement. It natively holds semi-structured JSON, and analytics tools query it directly without provisioning servers.

Why this answer

Amazon S3 is the correct choice because it provides highly durable, cost-effective object storage ideal for semi-structured JSON transaction logs that are written once and rarely accessed. S3's lifecycle policies can automatically transition such infrequently accessed data to S3 Glacier or S3 Glacier Deep Archive for even lower storage costs, making it the most economical option for this use case.

Exam trap

The trap here is that candidates may choose DynamoDB or Redshift because they support JSON natively, but they overlook the core requirement of cost-effective storage for rarely accessed data, which is best met by S3's low-cost object storage and lifecycle management features.

How to eliminate wrong answers

Option B (Amazon DynamoDB) is wrong because it is a NoSQL key-value and document database optimized for low-latency, high-throughput read/write operations, not for cost-effective archival storage of rarely accessed logs; storing large volumes of infrequently accessed JSON logs in DynamoDB would incur significant costs for provisioned throughput and storage. Option C (Amazon RDS) is wrong because it is a relational database service designed for transactional workloads with structured data and frequent queries, not for storing semi-structured JSON logs at low cost; it would require schema management and incur higher per-GB storage costs compared to S3. Option D (Amazon Redshift) is wrong because it is a petabyte-scale data warehouse optimized for complex analytical queries on structured and semi-structured data, not for simple, cost-effective archival storage; using Redshift for rarely accessed logs would be over-provisioned and expensive due to its compute and storage costs.

254
MCQmedium

A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources in Parquet format, and the schema evolves over time. Which approach allows querying the data with Amazon Athena while supporting schema evolution?

A.Use AWS Glue Data Catalog with crawlers to automatically update the table schema.
B.Define Hive-style partitions in Athena and manually update the schema.
C.Use S3 Select to query the data directly without a schema.
D.Use Amazon Redshift Spectrum with external tables and update the schema manually.
AnswerA

AWS Glue crawlers inspect Parquet data in Amazon S3 and populate the AWS Glue Data Catalog, which Athena queries. When new columns appear, re-running the crawler updates the table definition, satisfying the schema evolution requirement without manual DDL.

Why this answer

AWS Glue Data Catalog with crawlers automatically infers and updates the table schema as new Parquet files with evolving schemas are ingested into S3. This allows Athena to query the data using the latest schema without manual intervention, making it the ideal solution for schema evolution in a data lake.

Exam trap

The trap here is that candidates may think S3 Select or Redshift Spectrum can handle schema evolution automatically, but they lack the schema inference and versioning capabilities that AWS Glue Data Catalog provides for Athena.

How to eliminate wrong answers

Option B is wrong because manually updating the schema in Athena is error-prone and does not scale with frequent schema changes; Hive-style partitions alone do not handle schema evolution. Option C is wrong because S3 Select operates on individual objects and returns data in CSV/JSON format, not Parquet, and it does not support schema evolution or table-level queries across multiple files. Option D is wrong because Redshift Spectrum requires manual schema updates for external tables and is not designed for automatic schema evolution like AWS Glue Data Catalog.

255
MCQhard

A company runs an Apache Spark job on Amazon EMR that writes output to an S3 bucket. The job fails with the error 'S3AccessDeniedException' when writing the final output, but earlier stages succeed. The EMR cluster uses a service role and an instance profile. The S3 bucket policy allows access from the VPC only. What is the MOST likely cause?

A.The S3 bucket uses SSE-C encryption, and the EMR cluster does not have the encryption key.
B.The EMR service role does not have permissions to write to the S3 bucket.
C.The EMR cluster is not using a VPC endpoint for S3, so requests are denied by the bucket policy's VPC condition.
D.The S3 bucket is configured with 'Bucket owner enforced' setting for ACLs, and the EMR cluster's account is not the bucket owner.
AnswerC

The bucket policy restricts access to VPC, but since the Spark job runs on EMR, its requests originate from inside the VPC only if a VPC endpoint is used; otherwise, they come from public IPs.

Why this answer

The bucket policy restricts access to requests originating from the VPC, typically using a condition like `aws:SourceVpc`. If the EMR cluster does not use a VPC endpoint for S3 (either Gateway or Interface endpoint), traffic from the cluster to S3 traverses the public internet and does not match the VPC condition, causing the `S3AccessDeniedException`. Earlier stages may succeed if they use cached data or different paths, but the final write fails because it hits the bucket policy check.

Exam trap

The trap here is that candidates often assume the EMR service role (EMR_EC2_DefaultRole) is responsible for all S3 access, but in reality the instance profile (EC2 instance role) handles data plane operations, and the bucket policy's VPC condition is the key blocker when earlier stages succeed but final writes fail.

How to eliminate wrong answers

Option A is wrong because SSE-C encryption requires the client to provide the encryption key; if the key were missing, the error would be an encryption-related error (e.g., 'InvalidArgument' or 'AccessDenied' with a different message), not a generic 'S3AccessDeniedException'. Option B is wrong because the EMR service role is used for the cluster's service-level permissions (e.g., launching instances, reading logs), not for data access to S3; the instance profile (IAM role attached to EC2 instances) handles data read/write permissions, and the question states earlier stages succeed, indicating the instance profile has write permissions. Option D is wrong because the 'Bucket owner enforced' setting (S3 Object Ownership) controls ACLs and ownership of objects, not access permissions; it does not cause an 'S3AccessDeniedException' — it would affect who owns new objects, not whether the write is allowed.

256
MCQmedium

A company has an Amazon RDS for MySQL DB instance with read replicas. The primary DB instance fails. What is the correct procedure to promote a read replica to become the new primary?

A.Modify the read replica to be a Multi-AZ deployment and failover will occur.
B.RDS automatically fails over to the read replica within 5 minutes.
C.Manually promote the read replica to a standalone DB instance.
D.Delete the primary and the read replica will automatically become the primary.
AnswerC

Promoting a read replica manually converts it into a standalone DB instance, which is the only supported route when the primary fails without Multi-AZ. The stem specifies RDS for MySQL with read replicas, so no automatic failover exists; you must trigger promotion yourself to restore write capability.

Why this answer

When an Amazon RDS for MySQL primary DB instance fails, read replicas do not automatically become the new primary. The correct procedure is to manually promote the read replica using the AWS Management Console, CLI, or API, which converts it into a standalone DB instance. After promotion, you must update your application endpoints to point to the new primary, as RDS does not handle this automatically.

Exam trap

The trap here is that candidates confuse read replicas with Multi-AZ standby instances, assuming automatic failover applies to both, but RDS read replicas require manual promotion and do not provide automatic failover.

How to eliminate wrong answers

Option A is wrong because modifying a read replica to be Multi-AZ does not trigger a failover; Multi-AZ is a separate feature for high availability within a single region, and read replicas are not part of the Multi-AZ failover mechanism. Option B is wrong because RDS does not automatically fail over to a read replica; automatic failover only occurs with Multi-AZ deployments, not with read replicas. Option D is wrong because deleting the primary DB instance does not cause the read replica to automatically become the primary; the read replica remains a read-only copy until manually promoted.

257
MCQmedium

A data engineering team is using Amazon EMR to process large datasets stored in Amazon S3. The cluster uses Spot Instances for cost savings. During processing, the team notices that tasks are failing due to Spot Instance interruptions. The team needs to make the EMR job resilient to Spot interruptions without increasing costs significantly. Which solution should they implement?

A.Use EMR instance fleets with a mix of Spot and On-Demand, but set the allocation strategy to 'lowest price'.
B.Increase the number of core nodes using On-Demand instances.
C.Use only Spot Instances but enable automatic termination and checkpointing.
D.Use EMR instance fleets with a mix of Spot and On-Demand, setting the allocation strategy to 'diversified' and using On-Demand for core nodes.
AnswerD

Instance fleets with a diversified allocation strategy spread Spot capacity across multiple instance types and Availability Zones, reducing simultaneous interruption risk, while On-Demand core nodes preserve HDFS and task durability. This keeps costs low without losing resilience.

Why this answer

Using EMR instance fleets with a mix of Spot and On-Demand instances, setting the allocation strategy to 'diversified' and using On-Demand for core nodes, provides resilience against Spot interruptions while controlling costs. The 'diversified' strategy spreads Spot instances across multiple instance pools and Availability Zones, reducing the impact of a single Spot interruption. Core nodes, which handle HDFS data and run tasks, are kept on On-Demand to ensure stability, while task nodes can use Spot for cost savings.

Exam trap

The trap is assuming that any mix of Spot and On-Demand is sufficient, but the allocation strategy and node type assignment are critical: using 'lowest price' or putting core nodes on Spot can lead to frequent job failures despite having On-Demand instances.

How to eliminate wrong answers

Option A is wrong because the 'lowest price' allocation strategy concentrates Spot instances in the cheapest pool, which increases the risk of widespread interruptions if that pool's Spot capacity is reclaimed. Option B is wrong because increasing core nodes with On-Demand instances significantly raises costs and does not address the resilience of task nodes that may still be on Spot. Option C is wrong because using only Spot Instances, even with checkpointing, does not prevent job failures from interruptions; checkpointing helps resume but does not eliminate the need for On-Demand capacity for critical nodes.

Option D is correct because it balances cost and resilience by diversifying Spot instances and using On-Demand for core nodes.

258
MCQmedium

A company is migrating an on-premises Apache Cassandra database to Amazon Keyspaces. The database has a table with a partition key of 'user_id' and a clustering column of 'timestamp'. The application frequently queries the last 10 records for a given user. Which table design in Keyspaces would provide the BEST query performance for this access pattern?

A.Partition key: random column, clustering column: none.
B.Partition key: timestamp, clustering column: user_id.
C.Partition key: user_id, clustering column: none.
D.Partition key: user_id, clustering column: timestamp (descending order).
AnswerD

Descending clustering order stores rows on disk newest-first, so retrieving the last 10 records for a user reads one contiguous slice from the start of the partition rather than scanning the entire partition and sorting. This directly satisfies the frequent "last 10 records per user" access pattern with a single efficient query.

Why this answer

It preserves the original Cassandra table design with 'user_id' as the partition key and 'timestamp' as the clustering column in descending order. This allows Keyspaces to efficiently retrieve the last 10 records for a given user by performing a range query on the clustering column within a single partition, avoiding full table scans or cross-partition queries.

Exam trap

The trap here is that candidates may think a random partition key (Option A) or timestamp-based partition key (Option B) improves write distribution, but they overlook that the query pattern requires efficient reads within a single partition, which is best achieved by using the query filter column as the partition key and the sort column as the clustering key with the appropriate order.

How to eliminate wrong answers

Option A is wrong because using a random partition key with no clustering column would scatter data across partitions, requiring a full scan to find records for a specific user, which is highly inefficient. Option B is wrong because using 'timestamp' as the partition key would place each timestamp in a separate partition, making it impossible to query all records for a user without scanning multiple partitions, and the clustering column 'user_id' would not help retrieve the last 10 records per user efficiently. Option C is wrong because while 'user_id' as the partition key correctly groups data by user, having no clustering column means you cannot order records by timestamp, so retrieving the last 10 records would require fetching all records for that user and sorting them in application code, which is suboptimal.

259
MCQmedium

A data engineer sees this AWS Glue table definition in the Data Catalog. The engineer wants to query this table with Amazon Athena, but the query returns zero rows. What is the MOST likely cause?

A.The data files are not in the specified S3 location.
B.The SerDe library is incorrect for CSV files.
C.The table format CSV is not supported by Athena.
D.Athena cannot read tables from the Glue Data Catalog.
AnswerA

Athena reads the table's declared S3 location from the Data Catalog; if no objects exist at that prefix, the query scans nothing and returns zero rows. The partition metadata and schema may be valid, but the underlying files must reside at the specified location.

Why this answer

The most likely cause is that the data files are not in the specified S3 location. When an AWS Glue table is defined in the Data Catalog, Athena reads the table's metadata (including the S3 location) and then attempts to read the underlying data files from that exact path. If the files are missing, misnamed, or in a different prefix, Athena returns zero rows because there is no data to scan.

This is a common misconfiguration when the S3 path in the table definition does not match the actual data storage.

Exam trap

The trap here is that candidates often assume the issue is with the SerDe or format compatibility, but the most common real-world cause is simply that the data files are not present at the specified S3 location, leading to zero rows returned.

How to eliminate wrong answers

Option B is wrong because the SerDe library is not incorrect for CSV files; Athena uses the LazySimpleSerDe by default for CSV, which is fully supported and does not cause zero rows. Option C is wrong because CSV is a widely supported table format in Athena, and Athena can query CSV files natively. Option D is wrong because Athena is designed to read tables from the Glue Data Catalog; in fact, Athena and Glue Data Catalog are tightly integrated, and this is a standard use case.

260
MCQmedium

A company uses Amazon S3 to store sensitive data. The security team wants to ensure that all objects uploaded to a specific S3 bucket are automatically encrypted at rest using server-side encryption with AWS KMS managed keys (SSE-KMS). Which bucket policy statement should be added to enforce this requirement?

A.Deny put requests where 's3:x-amz-server-side-encryption' is 'aws:kms'
B.Deny put requests where 's3:x-amz-server-side-encryption' is not 'aws:kms'
C.Deny put requests where 's3:x-amz-server-side-encryption' is not 'AES256'
D.Deny put requests where 's3:x-amz-server-side-encryption' is not set
AnswerB

A bucket policy denying `s3:PutObject` when `s3:x-amz-server-side-encryption` is absent or not `aws:kms` enforces SSE-KMS at upload time, satisfying the requirement that every object be encrypted with AWS KMS managed keys. The condition key inspects the request header, blocking non-compliant uploads before they succeed.

Why this answer

It denies any S3 PUT request that does not include the `x-amz-server-side-encryption` header set to `aws:kms`, thereby enforcing SSE-KMS encryption for all objects uploaded to the bucket. This bucket policy condition ensures that only requests specifying AWS KMS-managed keys are allowed, meeting the security team's requirement for automatic encryption at rest with SSE-KMS.

Exam trap

The DEA-C01 exam often tests the distinction between enforcing a specific encryption type (SSE-KMS) versus simply requiring encryption (any type), so candidates may incorrectly choose Option D (deny if not set) or Option C (deny if not AES256) because they confuse 'encryption at rest' with 'SSE-KMS specifically'.

How to eliminate wrong answers

Option A is wrong because it denies PUT requests where `s3:x-amz-server-side-encryption` is `aws:kms`, which would block the very encryption method required, making it impossible to upload objects with SSE-KMS. Option C is wrong because it denies PUT requests where encryption is not `AES256`, which would enforce SSE-S3 (AES256) instead of SSE-KMS, failing the requirement for KMS-managed keys. Option D is wrong because it denies PUT requests where the encryption header is not set, which would block unencrypted uploads but does not specifically enforce SSE-KMS; it would also allow SSE-S3 or other encryption types if the header is present, missing the specific KMS requirement.

261
MCQeasy

A company is using Amazon DynamoDB for a gaming application. They want to store player session data that expires after 24 hours. Which DynamoDB feature should be used?

A.Time to Live (TTL)
B.DynamoDB Streams
C.Global Tables
D.Point-in-Time Recovery
AnswerA

Time to Live (TTL) lets you define an epoch-timestamp attribute so DynamoDB automatically deletes expired items at no write cost, satisfying the 24-hour session expiry without custom cleanup jobs or scans. Deletion typically occurs within 48 hours of expiry, though expired items are already hidden from reads.

Why this answer

Amazon DynamoDB Time to Live (TTL) allows you to define a per-item timestamp attribute that automatically deletes items after a specified duration. For the gaming session data that must expire after 24 hours, you can set the TTL attribute to the current time plus 24 hours, and DynamoDB will asynchronously delete expired items without any additional cost or write operations.

Exam trap

The trap here is that candidates may confuse DynamoDB Streams (which can react to deletions) with the actual mechanism that performs the deletion, or assume Point-in-Time Recovery can be used to 'roll back' expired data, neither of which addresses automatic expiration.

How to eliminate wrong answers

Option B (DynamoDB Streams) is wrong because it captures a time-ordered sequence of item-level changes (inserts, updates, deletes) in a DynamoDB table, but it does not automatically expire or delete data; it is used for event-driven processing or replication, not for scheduled data removal. Option C (Global Tables) is wrong because it provides multi-region, fully replicated tables for low-latency access and disaster recovery, but it has no built-in mechanism to expire or delete items based on time. Option D (Point-in-Time Recovery) is wrong because it enables continuous backups of DynamoDB table data to restore to any point within the last 35 days, but it does not delete or manage the lifecycle of individual items.

262
Multi-Selectmedium

Which TWO actions can reduce the cost of an Amazon S3 bucket that stores infrequently accessed data? (Choose 2.)

Select 2 answers
A.Enable cross-region replication
B.Enable versioning to keep multiple versions
C.Use lifecycle policies to expire objects after a certain period
D.Enable MFA Delete for extra security
E.Transition objects to S3 Standard-IA after 30 days
AnswersC, E

Expiration deletes unneeded objects.

Why this answer

Lifecycle policies allow you to define rules that automatically expire (delete) objects after a specified period, which directly reduces storage costs by removing data that is no longer needed. For infrequently accessed data, deleting obsolete objects prevents paying for unnecessary storage over time.

Exam trap

The DEA-C01 exam often tests the misconception that enabling versioning or replication reduces costs, when in fact both increase storage and transfer costs, while lifecycle policies and storage class transitions are the correct cost-saving mechanisms.

263
Multi-Selecteasy

A data engineer is setting up Amazon S3 event notifications to trigger an AWS Lambda function when new objects are uploaded. Which TWO actions are required to enable this?

Select 2 answers
A.Add a resource-based policy to the Lambda function to allow S3 to invoke it.
B.Enable S3 versioning on the bucket.
C.Create an S3 bucket policy that grants S3 permission to invoke Lambda.
D.Configure an event notification on the S3 bucket for s3:ObjectCreated:* events.
E.Set up an Amazon CloudWatch Events rule to detect S3 uploads.
AnswersA, D

S3 invokes Lambda through a resource-based policy, so the function must grant s3.amazonaws.com permission via AddPermission, scoped to the bucket's source ARN. This satisfies the stem's requirement that S3 be authorised to trigger the function, without which the notification configuration fails with an access error.

Why this answer

Lambda functions use a resource-based policy (also known as a function policy) to grant permissions to other AWS services, such as S3, to invoke the function. Without this policy, S3 will receive an access denied error when trying to trigger the Lambda function. Option D is correct because you must configure an S3 event notification on the bucket for the `s3:ObjectCreated:*` event type to instruct S3 to send a notification to the Lambda function when new objects are uploaded.

Exam trap

The trap here is that candidates often think an S3 bucket policy is needed to allow S3 to invoke Lambda, but in reality, the permission must be granted on the Lambda function's resource-based policy, not on the bucket.

264
MCQeasy

A data engineer must choose a storage service for a new application that requires single-digit millisecond latency at any scale, a flexible schema, and automatic scaling of throughput without provisioning capacity. The access pattern is key-value lookups by user ID with occasional range queries on a sort key. Which AWS service should the engineer select?

A.Amazon RDS for PostgreSQL with automatic storage scaling
B.Amazon S3 with S3 Select
C.Amazon Redshift with a distribution key on user ID
D.Amazon DynamoDB with on-demand capacity mode
AnswerD

DynamoDB delivers consistent single-digit millisecond latency at any scale, supports a flexible schema with a partition key and optional sort key, and on-demand capacity mode scales read and write throughput automatically without provisioning. Key-value lookups by user ID map to the partition key, and range queries map to the sort key, matching the access pattern exactly.

Why this answer

DynamoDB is purpose-built for key-value and document workloads needing consistent single-digit millisecond latency at any scale. A partition key on user ID serves point lookups, and a sort key supports efficient range queries. On-demand capacity mode removes provisioning and scales automatically with traffic, while the flexible schema accommodates evolving attributes without migrations, matching every stated requirement.

Exam trap

The trap here is assuming a relational database such as Amazon RDS or a warehouse such as Redshift can deliver single-digit millisecond latency at any scale, when that guarantee belongs to DynamoDB's key-value design.

265
MCQhard

A data engineer manages an Amazon Redshift cluster that experiences performance degradation during complex analytical queries. The engineer notices that some queries spill to disk. The engineer wants to improve query performance by optimizing the distribution style and sort keys. Which action should the engineer take first?

A.Change all tables to use EVEN distribution style to balance data across nodes.
B.Analyze the query execution plans to identify tables with high redistribution or broadcast steps.
C.Set the sort key to the column most frequently used in WHERE clauses for all tables.
D.Increase the number of nodes in the cluster to add more memory and storage.
AnswerB

Query execution plans reveal how data is distributed and joined. High redistribution or broadcast steps indicate suboptimal distribution styles, causing data movement across nodes. By analyzing these plans, the engineer can identify which tables need distribution key changes. This is the first step before making changes, as it provides data-driven insights. It ensures that modifications target the actual bottlenecks.

Why this answer

Analyzing query execution plans is the essential first step to identify performance bottlenecks such as data redistribution or broadcast operations. These indicate suboptimal distribution styles. Once identified, the engineer can make informed changes to distribution keys and sort keys.

The other options are either premature or not targeted at the root cause.

Exam trap

The trap here is jumping to schema changes or scaling without first diagnosing the actual query execution bottlenecks.

266
MCQmedium

A data engineer manages an Amazon Redshift cluster that runs a nightly ETL load followed by complex analytical queries. Users report that queries during the day are slower than expected, and the team wants to isolate the ETL workload so it cannot consume resources needed by the analytical queries. The cluster uses provisioned nodes. What is the MOST appropriate solution?

A.Increase the number of nodes in the cluster to provide more resources for all workloads.
B.Enable concurrency scaling on the cluster and route all queries through it.
C.Create a separate Redshift workload management (WLM) queue for the ETL role and assign the analytical queries to a different queue.
D.Move the ETL workload to a separate Redshift cluster and use Amazon Redshift Spectrum for analytics.
AnswerC

Redshift WLM lets you define multiple queues with dedicated memory and concurrency slots, and you can route queries to queues based on user groups or query groups. Assigning ETL to its own queue prevents it from starving the analytics queue. This directly isolates the workloads on the same provisioned cluster without extra infrastructure.

Why this answer

Redshift WLM provides workload isolation by letting you define multiple queues, each with its own memory allocation and query concurrency, and by routing queries to queues based on user groups or query groups. Placing the ETL job in a dedicated queue ensures it cannot consume the memory and slots reserved for analytical queries, which addresses the slowdowns without adding clusters or nodes.

Exam trap

The trap here is assuming that adding nodes or enabling concurrency scaling will isolate workloads, when isolation requires separate WLM queues with defined memory and concurrency.

267
MCQmedium

A media company stores video files in an S3 bucket. The files are processed by a fleet of EC2 instances that read the files, add watermarks, and write the output back to the same bucket. Recently, the processing jobs have been failing with '500 Internal Server Error' and '503 Slow Down' errors. The data engineer checks the S3 bucket metrics and sees that the PUT/GET request rate is consistently above 5,500 requests per second for a single prefix. The engineer needs to resolve the errors with minimal changes to the application code. Which course of action should the engineer take?

A.Use S3 Batch Operations to process the files.
B.Increase the number of EC2 instances to process files in parallel.
C.Enable S3 Transfer Acceleration on the bucket to improve throughput.
D.Modify the application to add a random hash prefix to the object keys to distribute load across multiple prefixes.
AnswerD

S3 partitions request throughput per prefix, so exceeding roughly 3,500 PUT or 5,500 GET requests per second on one prefix triggers 503 Slow Down throttling. Adding a random hash prefix spreads keys across many partitions, restoring throughput without changing the bucket or application logic significantly.

Why this answer

S3 automatically partitions request traffic by key prefix, and each partition supports roughly 3,500 PUT/COPY/POST/DELETE or 5,500 GET/HEAD requests per second. When a single prefix exceeds these limits, S3 returns 503 Slow Down errors. Adding a random hash prefix to object keys distributes the load across many prefixes, each with its own request rate limit, thereby eliminating throttling without changing the core application logic.

This is the recommended best practice for high-throughput workloads.

Exam trap

DEA-C01 often tests the misconception that adding more compute resources (EC2 instances) or enabling Transfer Acceleration will solve S3 throttling, when the actual fix is to distribute the request load across multiple key prefixes.

How to eliminate wrong answers

Option A is wrong because S3 Batch Operations is designed for bulk operations on existing objects (like copying or tagging) and does not increase request rate limits or resolve throttling for ongoing PUT/GET operations. Option B is wrong because adding more EC2 instances would increase the request rate further, exacerbating the throttling issue rather than alleviating it. Option C is wrong because S3 Transfer Acceleration speeds up transfers over long distances by using AWS edge locations, but it does not increase the per-prefix request rate limits and will not resolve 503 Slow Down errors caused by exceeding those limits.

268
MCQmedium

A data engineer is building a data lake on Amazon S3 and must enforce that all objects containing personally identifiable information are encrypted with a customer managed AWS KMS key, while allowing automatic key rotation and audit of key usage. Objects must remain readable by an AWS Glue job and an Amazon Athena workgroup. Which configuration should the engineer choose?

A.Client-side encryption with an application managed key stored in the application code
B.SSE-KMS with the AWS managed key aws/s3 and default bucket encryption
C.SSE-KMS with a customer managed key, key rotation enabled, and bucket policy enforcing the key
D.SSE-S3 with bucket versioning enabled
AnswerC

SSE-KMS with a customer managed key gives the organization control over the key policy, enables automatic annual rotation, and records every cryptographic operation in AWS CloudTrail for audit. A bucket policy that denies uploads not using the specified key enforces consistency. Granting the Glue role and Athena workgroup kms:Decrypt and kms:GenerateDataKey permissions allows both services to read and write objects.

Why this answer

A customer managed KMS key with SSE-KMS satisfies control, rotation, and auditability: the key policy scopes who may use it, automatic rotation can be enabled, and CloudTrail logs every Encrypt, Decrypt, and GenerateDataKey call. A bucket policy that denies noncompliant uploads enforces the standard, and granting the Glue role and Athena workgroup decrypt permissions keeps analytics working. AWS managed keys and SSE-S3 cannot meet the customer control and rotation requirements.

Exam trap

The trap here is treating any KMS-based encryption as equivalent, when only a customer managed key provides configurable rotation and a key policy that can be audited and restricted.

269
MCQmedium

A company uses Amazon Redshift for data warehousing. The data engineering team notices that queries are slow due to high disk I/O. The team wants to improve query performance without changing the cluster configuration. Which action should the team take?

A.Increase the number of nodes in the cluster.
B.Redesign tables with appropriate sort keys and distribution styles.
C.Run the ANALYZE command to update table statistics.
D.Run the VACUUM command to reclaim disk space.
AnswerB

Sort keys reduce the blocks scanned by enabling zone-map pruning, while distribution styles co-locate joined rows on the same slice, cutting inter-node network traffic. Both target disk I/O and skew directly, satisfying the constraint of improving performance without altering cluster configuration.

Why this answer

Redesigning tables with appropriate sort keys and distribution styles directly addresses high disk I/O by minimizing data scanning and reducing data movement across nodes. Sort keys enable Redshift to skip irrelevant blocks via zone maps, while distribution styles (KEY, ALL, EVEN) optimize data locality for joins and aggregations, reducing I/O without changing cluster configuration.

Exam trap

The trap here is that candidates confuse maintenance commands (ANALYZE, VACUUM) with design changes, or think scaling out (adding nodes) is allowed when the question explicitly forbids changing cluster configuration.

How to eliminate wrong answers

Option A is wrong because increasing the number of nodes changes the cluster configuration, which the question explicitly prohibits. Option C is wrong because ANALYZE updates table statistics for the query planner but does not reduce disk I/O caused by poor data layout or data movement. Option D is wrong because VACUUM reclaims disk space from deleted rows and sorts data, but it does not fundamentally redesign tables to reduce I/O; it only maintains existing design.

270
MCQhard

A company uses Amazon S3 to store large datasets. The data engineering team needs to provide access to specific objects in the bucket to external partners using presigned URLs. Each URL should expire after 12 hours. The team wants to ensure that the presigned URLs cannot be used to access other objects in the bucket. Which approach should be taken?

A.Create an IAM role for each partner and attach a policy that grants access to specific objects.
B.Generate presigned URLs using the AWS SDK, specifying the exact object key and expiration time.
C.Use a bucket policy that allows access only from the partner's IP address range.
D.Use CloudFront signed URLs with a custom policy that restricts access to specific objects.
AnswerB

Generating presigned URLs with the AWS SDK while specifying the exact object key scopes each URL to a single object, and the 12-hour expiry parameter satisfies the required lifetime. This prevents access to other objects in the bucket.

Why this answer

Presigned URLs generated via the AWS SDK allow you to specify the exact object key and expiration time, ensuring that the URL grants access only to that specific object for the defined 12-hour period. This approach uses the secret key of the IAM user or role to sign the URL, and the signature is tied to the object key, so the URL cannot be reused to access other objects in the bucket.

Exam trap

The DEA-C01 exam often tests the distinction between presigned URLs (which are tied to a specific object key and expiration) and bucket policies or IAM roles (which grant broader access), leading candidates to overcomplicate the solution with CloudFront or IP-based restrictions when a simple SDK-generated presigned URL is sufficient.

How to eliminate wrong answers

Option A is wrong because creating an IAM role for each partner and attaching a policy that grants access to specific objects does not inherently enforce time-limited access; the role would need additional mechanisms like STS to generate temporary credentials, and it does not provide the simplicity of a single URL. Option C is wrong because a bucket policy that allows access only from the partner's IP address range would grant access to all objects in the bucket (or a broader set) rather than restricting to specific objects, and it does not provide time-limited access. Option D is wrong because CloudFront signed URLs require CloudFront distribution and custom origin setup, which adds unnecessary complexity and cost; while they can restrict access to specific objects, they are not the simplest or most direct solution for S3 presigned URLs, and the question specifically asks for presigned URLs.

271
MCQeasy

A data engineer needs to store JSON documents that are accessed by a serverless application using AWS Lambda. The documents are frequently updated and need low latency (single-digit milliseconds) for read and write operations. Which AWS service should the engineer use?

A.Amazon DynamoDB
B.Amazon ElastiCache for Redis
C.Amazon S3 (with S3 Select)
D.Amazon RDS for MySQL
AnswerA

DynamoDB delivers consistent single-digit-millisecond latency for both reads and writes at any scale, and its serverless, fully managed design pairs directly with Lambda. The stem's frequent updates and low-latency requirement rule out S3, which offers higher and more variable latency.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value and document database that provides single-digit millisecond latency for read and write operations at any scale. It natively supports JSON documents, integrates directly with AWS Lambda via the AWS SDK, and handles frequent updates efficiently through its auto-scaling and on-demand capacity modes, making it ideal for serverless applications requiring low-latency data access.

Exam trap

The trap here is that candidates often confuse ElastiCache for Redis as a primary data store due to its low latency, overlooking that it is an in-memory cache with no built-in persistence guarantees, whereas DynamoDB provides both low latency and durable, persistent storage for JSON documents.

How to eliminate wrong answers

Option B is wrong because Amazon ElastiCache for Redis is an in-memory cache, not a durable data store; while it offers sub-millisecond latency, it is typically used for caching or session management and requires a separate persistent database to avoid data loss on node failure, making it unsuitable as the primary store for frequently updated JSON documents that must persist. Option C is wrong because Amazon S3 is an object storage service with eventual consistency for overwrite PUTS and higher latency (typically tens to hundreds of milliseconds) for read operations, and S3 Select is a server-side filtering feature that does not reduce latency for individual document reads or writes; it is not designed for frequent, low-latency updates. Option D is wrong because Amazon RDS for MySQL is a relational database that requires schema definition, does not natively store JSON as a first-class document model (though it supports JSON data type, it lacks the flexible schema and single-digit millisecond read/write performance of DynamoDB for key-value access patterns), and incurs higher operational overhead for scaling and connection management in a serverless architecture.

272
Multi-Selectmedium

A company uses Amazon Redshift for analytics. The data engineering team wants to improve query performance for frequently used aggregate queries. Which TWO actions would help achieve this?

Select 2 answers
A.Increase the number of WLM query queues
B.Use distribution keys to collocate data on the same node slices
C.Run the VACUUM command to reclaim space from deleted rows
D.Define appropriate sort keys on the tables
E.Increase the number of nodes in the cluster
AnswersB, D

Distribution keys determine which node slice stores each row, so collocating joined or aggregated rows on the same slice lets Redshift perform local joins and partial aggregation, cutting data movement across the network during aggregate queries.

Why this answer

Option B is correct because choosing an appropriate distribution key collocates matching rows on the same node slices, so joins and aggregations can be processed locally without expensive data redistribution (broadcast or shuffle) across the cluster, directly speeding up frequently used aggregate queries. Option D is correct because defining appropriate sort keys physically orders data on disk by the key columns, enabling zone maps to skip irrelevant blocks and allowing efficient range-restricted scans and merge joins, which reduces the data read for aggregate queries. Option A is not correct because adding WLM query queues only changes concurrency and memory allocation among query groups; it does not by itself make an individual aggregate query faster.

Option C is not correct because VACUUM reclaims space from deleted rows and re-sorts data, which is a maintenance operation rather than a design change that improves aggregate query performance. Option E is not correct because adding nodes increases cluster capacity and parallelism but does not address the underlying data layout, so poorly distributed or unsorted tables can still cause slow aggregate queries.

Exam trap

The trap here is that candidates often confuse VACUUM (which reclaims space) with performance optimization for queries, or assume adding nodes always improves query speed without considering the overhead of data redistribution.

273
MCQmedium

A data engineer is using Amazon Redshift and needs to improve query performance for a large fact table that is frequently joined with a much smaller dimension table. The engineer wants to minimize data movement during joins. Which distribution style should be used for the dimension table?

A.DISTSTYLE AUTO
B.DISTSTYLE KEY
C.DISTSTYLE ALL
D.DISTSTYLE EVEN
AnswerC

DISTSTYLE ALL replicates the entire dimension table to every compute node. When joining with a large fact table, each node already has a copy of the dimension table, eliminating the need to redistribute the dimension data during the join. This minimizes data movement and improves join performance, which is ideal for small dimension tables.

Why this answer

For a small dimension table that is frequently joined with a large fact table, DISTSTYLE ALL replicates the dimension table to all nodes. This ensures that each node has a local copy, eliminating the need to broadcast or redistribute the dimension data during joins. As a result, data movement is minimized, and join performance is significantly improved.

Exam trap

The trap here is assuming that automatic distribution or distributing on the join key is always best, when for small dimension tables, replicating the entire table with DISTSTYLE ALL is often more efficient.

274
MCQhard

A data engineer notices that an Amazon Redshift cluster’s storage usage is increasing rapidly due to many UPDATE and DELETE operations. The engineer needs to reclaim storage space and improve query performance. Which action should be taken?

A.Run VACUUM command
B.UNLOAD the table to S3 and reload
C.Increase cluster node count
D.Run ANALYZE command
AnswerA

VACUUM reclaims space from deleted rows and re-sorts unsorted regions, restoring sequential scan performance degraded by frequent UPDATE and DELETE operations. This directly addresses the stem's storage growth and query-performance constraints on the Redshift cluster.

Why this answer

The VACUUM command in Amazon Redshift reclaims disk space occupied by deleted or updated rows and re-sorts the data according to the table's sort keys. This directly addresses the storage increase from UPDATE/DELETE operations and improves query performance by restoring the physical order of rows, which reduces the number of blocks scanned.

Exam trap

The trap here is that candidates confuse ANALYZE with VACUUM, thinking updating statistics will also reclaim storage, when in fact ANALYZE only refreshes metadata for the query optimizer and has no effect on physical storage.

How to eliminate wrong answers

Option B is wrong because unloading the table to S3 and reloading is a heavy, manual process that does not reclaim space in place and can be avoided with a simple VACUUM; it also incurs additional S3 costs and time. Option C is wrong because increasing the cluster node count adds more storage and compute capacity but does not reclaim the existing wasted space from deleted rows, and it may not improve performance if the underlying data is fragmented. Option D is wrong because the ANALYZE command only updates table statistics for the query planner, it does not reclaim storage space or physically reorganize data affected by UPDATE/DELETE operations.

275
MCQeasy

A data engineer needs to store semi-structured data (JSON logs) from thousands of IoT devices. The data must be schema-less, highly scalable, and support low-latency queries by device ID and timestamp. Which AWS service should the engineer use?

A.Amazon RDS for PostgreSQL
B.Amazon Redshift
C.Amazon DynamoDB
D.Amazon S3
AnswerC

DynamoDB stores JSON as native map and list attributes without a fixed schema, scales horizontally, and a composite partition key of device ID plus sort key on timestamp delivers low-latency item queries — matching the schema-less, scalable, low-latency constraints exactly.

Why this answer

Amazon DynamoDB is the correct choice because it is a fully managed NoSQL key-value and document database that natively supports semi-structured JSON data, schema-less design, and automatic scaling. Its partition key (device ID) and sort key (timestamp) enable low-latency, single-millisecond queries by device ID and timestamp, making it ideal for high-throughput IoT log ingestion.

Exam trap

The trap here is that candidates often confuse Amazon S3's ability to store JSON files with the ability to query them efficiently, overlooking that S3 lacks native indexing and low-latency query support, which DynamoDB provides through its key-value access pattern.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for PostgreSQL is a relational database with a fixed schema, requiring predefined tables and indexes for JSON data, which cannot handle schema-less IoT logs at scale without manual sharding or performance tuning. Option B is wrong because Amazon Redshift is a columnar data warehouse optimized for analytical queries on structured data, not for low-latency point queries by device ID and timestamp, and its schema-on-write model conflicts with schema-less requirements. Option D is wrong because Amazon S3 is an object store that can store JSON logs but lacks native indexing and low-latency query capabilities; querying by device ID and timestamp would require scanning or external services like Athena, adding latency and complexity.

276
MCQeasy

A company stores application logs in an Amazon S3 bucket. A compliance policy states that log objects must be retained for exactly 90 days and then permanently deleted, and that no one, including administrators, should be able to delete them earlier. The data engineer must enforce this with the least effort. What should the engineer do?

A.Create an S3 Lifecycle rule that transitions objects to S3 Glacier Deep Archive after 90 days.
B.Apply an S3 Object Lock retention period of 90 days in compliance mode to the bucket, and configure a lifecycle rule to expire objects after 90 days.
C.Use AWS Backup to create a vault with a 90-day retention and a vault lock in compliance mode.
D.Enable S3 Versioning and add a bucket policy that denies s3:DeleteObject to all principals.
AnswerB

S3 Object Lock in compliance mode prevents any user, including the root user, from deleting or overwriting an object version until the retention period expires. Setting a 90-day retention plus a lifecycle expiration rule enforces both the immutability and the automatic deletion after 90 days, satisfying the policy with minimal ongoing effort.

Why this answer

S3 Object Lock in compliance mode provides WORM protection that even the root user cannot override before the retention period ends, which satisfies the immutability requirement. Pairing it with a lifecycle rule that expires objects after 90 days ensures automatic permanent deletion at the required time. Together they enforce the compliance policy with minimal operational effort.

Exam trap

The trap here is treating versioning or lifecycle transitions as immutability controls, when only Object Lock in compliance mode prevents deletion by any principal during the retention period.

277
MCQhard

A data engineer is configuring an Amazon S3 bucket to store sensitive financial data. The company requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption key be automatically rotated every year. The engineer creates a KMS customer managed key and enables automatic rotation. When uploading objects using the AWS CLI, the engineer uses the --sse aws:kms parameter but does not specify a key ID. What is the result of this configuration?

A.The objects are encrypted with the customer managed key because it is the only KMS key in the account.
B.The objects are encrypted with the customer managed key, and automatic rotation applies because the key is the default KMS key for the account.
C.The objects are encrypted with the AWS managed key for S3 (aws/s3), not the customer managed key.
D.The upload fails because a KMS key ID is required when using --sse aws:kms.
AnswerC

When you specify --sse aws:kms without a key ID, S3 uses the AWS managed key for S3 (aws/s3) by default. This key is managed by AWS and does not support automatic rotation configuration by the customer. To use the customer managed key, the engineer must specify its key ARN or alias with --sse-kms-key-id.

Why this answer

To use a specific KMS customer managed key for S3 default encryption or per-object encryption, you must explicitly provide the key ARN or alias. Omitting the key ID causes S3 to fall back to the AWS managed key (aws/s3), which does not meet the requirement of using a customer managed key with annual rotation.

Exam trap

The trap here is assuming that specifying --sse aws:kms automatically uses a customer managed key, when in fact it defaults to the AWS managed key unless a key ID is provided.

278
Multi-Selecthard

Which THREE factors should be considered when choosing a partition key for an Amazon DynamoDB table?

Select 3 answers
A.The partition key should be chosen to maximize the size of items in each partition.
B.If the table has a write-heavy workload, the partition key should distribute writes evenly.
C.The partition key should align with the most common query access pattern.
D.The partition key should be chosen to minimize read capacity unit consumption.
E.The partition key should have high cardinality to distribute data evenly.
AnswersB, C, E

Even write distribution across partitions prevents hot partitions, which throttle throughput when write volume is high. DynamoDB hashes the partition key to place items, so a high-cardinality key spreads writes evenly. This directly satisfies the stem's write-heavy workload constraint, sustaining provisioned capacity without request throttling.

Why this answer

Option B is correct because a write-heavy workload requires the partition key to spread write traffic across many partitions; otherwise a hot partition throttles throughput, since each partition supports a limited write capacity (up to 1,000 WCU per partition). Option C is correct because DynamoDB retrieves items by partition key, so aligning the key with the most common query access pattern enables efficient Query operations instead of expensive Scan operations. Option E is correct because high cardinality produces many distinct partition key values, which distributes items and traffic evenly across partitions and avoids hot partitions.

Option A is not a valid factor because item size does not determine partition key choice; DynamoDB partitions data by key value, and large items only consume more capacity, not improve partitioning. Option D is not a valid factor because RCU consumption is driven by item size and consistency model, not by selecting a partition key to minimize reads.

Exam trap

The trap here is that candidates may think maximizing item size (Option A) or minimizing RCU consumption (Option D) are primary factors, when in fact even distribution and access pattern alignment are the critical design principles for DynamoDB partition keys.

279
MCQmedium

A data engineer manages an Amazon S3 data lake with millions of small JSON files ingested continuously. Amazon Athena queries over this data are slow and expensive because each query scans many small objects. The engineer wants to improve query performance and reduce cost without changing the data content. Which solution should the engineer implement?

A.Use AWS Glue ETL to compact the small files into larger Parquet files partitioned by common query filters.
B.Increase the Athena query result reuse cache TTL to 7 days.
C.Convert the S3 bucket to S3 Intelligent-Tiering to improve read throughput.
D.Enable S3 Transfer Acceleration on the bucket to speed up Athena query reads.
AnswerA

Compacting small JSON files into larger Parquet files reduces the number of objects Athena must list and open, and Parquet's columnar format allows Athena to scan only needed columns. Partitioning by common filters further reduces data scanned. This directly addresses both performance and cost without altering the underlying data semantics, making it the most effective solution for this scenario.

Why this answer

The scenario describes slow and costly Athena queries caused by many small JSON files. Compacting into larger Parquet files and partitioning by common filters reduces the number of objects scanned and leverages columnar storage to scan less data. This directly improves performance and lowers cost without changing data content, making it the correct solution.

Exam trap

The trap here is assuming that S3 storage-class or transfer features can improve Athena query performance, when the real issue is file format and object count.

280
MCQeasy

A data engineer needs to store semi-structured JSON data from IoT devices. The data is written frequently and read occasionally. Which AWS service is MOST cost-effective for this use case?

A.Amazon ElastiCache for Redis
B.Amazon DynamoDB
C.Amazon RDS for MySQL
D.Amazon Redshift
AnswerB

DynamoDB is a serverless key-value store billed per request, so frequent writes and occasional reads incur no idle capacity cost, unlike always-on provisioned databases. It natively stores JSON documents, satisfying the semi-structured IoT payload requirement at the lowest cost.

Why this answer

Amazon DynamoDB is the most cost-effective choice because it is a fully managed NoSQL key-value and document database that natively supports semi-structured JSON data, offers single-digit millisecond latency for frequent writes, and provides a pay-per-request pricing model ideal for workloads with occasional reads. Its on-demand capacity mode automatically scales to handle high write throughput without provisioning, making it cheaper than provisioned alternatives for spiky or unpredictable IoT ingestion patterns.

Exam trap

The trap here is that candidates often choose Amazon ElastiCache for Redis due to its speed and JSON module support, but they overlook that it is not designed for durable, cost-effective long-term storage of semi-structured data, and DynamoDB's native JSON support and pay-per-request pricing make it the more economical choice for this specific write-frequent, read-occasional pattern.

How to eliminate wrong answers

Option A is wrong because Amazon ElastiCache for Redis is an in-memory cache designed for sub-millisecond read-heavy workloads and ephemeral data, not for durable storage of semi-structured JSON from IoT devices; it lacks native JSON document storage (though RedisJSON module exists, it adds cost and complexity) and is significantly more expensive per GB than DynamoDB for persistent data. Option C is wrong because Amazon RDS for MySQL is a relational database that requires schema definition, making it inefficient for semi-structured JSON data that varies in fields; it incurs higher costs due to provisioned IOPS and storage, and its write performance is limited by the underlying instance size and transaction overhead. Option D is wrong because Amazon Redshift is a columnar data warehouse optimized for complex analytical queries on large datasets, not for high-frequency writes from IoT devices; its minimum cost is high (starts at ~$0.25/hour for dc2.large), and it is overkill for occasional reads of semi-structured JSON, leading to wasted expenditure.

281
MCQmedium

A data engineer is using AWS Glue to process data stored in Amazon S3. The engineer needs to ensure that the AWS Glue job can access the S3 bucket securely without hardcoding credentials. Which approach should the engineer use?

A.Embed the AWS access key and secret key directly in the Glue job script.
B.Create an IAM role with the necessary S3 permissions and attach it to the AWS Glue job.
C.Store AWS credentials in AWS Secrets Manager and retrieve them within the Glue job script.
D.Use an Amazon S3 bucket policy that allows public read access to the data.
AnswerB

AWS Glue jobs assume an IAM role that grants permissions to access AWS resources. By attaching a role with the appropriate S3 permissions, the job can securely access the bucket without embedding credentials. This is the recommended best practice for AWS services. It also allows for fine-grained access control and auditing.

Why this answer

Attaching an IAM role with the necessary S3 permissions to the AWS Glue job is the secure and recommended way to grant access without hardcoding credentials. IAM roles provide temporary credentials and follow the principle of least privilege. The other options either introduce security risks or unnecessary complexity.

Exam trap

The trap here is considering Secrets Manager or hardcoded credentials when IAM roles are the native, secure method for service-to-service authentication.

282
MCQhard

A data engineer is designing a multi-region disaster recovery solution for Amazon RDS for PostgreSQL. The primary region must have a standby in a different Availability Zone, and the secondary region must have a readable replica that can be promoted in case of failure. Which configuration meets these requirements?

A.Use a single-AZ primary and enable automatic backups
B.Enable Multi-AZ in the primary region and create a cross-region read replica
C.Use a single-AZ primary and create a cross-region read replica
D.Enable Multi-AZ in both primary and secondary regions
AnswerB

Multi-AZ maintains a synchronous standby in a separate Availability Zone for automatic failover, while a cross-region read replica provides a readable copy in the secondary region that can be promoted. Together they satisfy both the in-region standby and cross-region readable replica constraints.

Why this answer

It meets both requirements: Multi-AZ in the primary region provides a synchronous standby in a different Availability Zone for high availability, and a cross-region read replica in the secondary region provides an asynchronous, readable copy that can be promoted to a standalone primary during a regional failure. This combination ensures both intra-region fault tolerance and inter-region disaster recovery.

Exam trap

The trap here is that candidates often confuse Multi-AZ (synchronous, for high availability within a region) with cross-region read replicas (asynchronous, for disaster recovery), and may incorrectly assume that Multi-AZ alone provides cross-region failover or that a single-AZ primary with a read replica satisfies the intra-region standby requirement.

How to eliminate wrong answers

Option A is wrong because a single-AZ primary with automatic backups does not provide a standby in a different Availability Zone, nor does it create a readable replica in a secondary region; backups are for point-in-time recovery, not for immediate failover or read scaling. Option C is wrong because a single-AZ primary lacks the required standby in a different Availability Zone within the primary region; the cross-region read replica only addresses the secondary region requirement. Option D is wrong because enabling Multi-AZ in both regions does not create a cross-region read replica; Multi-AZ in the secondary region provides a standby within that region but does not establish a readable replica that can be promoted from the primary region.

283
Multi-Selectmedium

A company uses Amazon DynamoDB for a gaming application. The application experiences throttling during peak hours. The table's read and write capacity is provisioned. Which TWO actions can reduce throttling?

Select 2 answers
A.Enable TTL (time to live) on the table to automatically delete old items
B.Enable DynamoDB auto scaling for the table
C.Increase the provisioned read capacity units (RCUs)
D.Implement DynamoDB Accelerator (DAX) to cache read requests
E.Add a DynamoDB Global Table for the table
AnswersB, D

Auto scaling adjusts provisioned capacity based on traffic.

Why this answer

DynamoDB auto scaling (Option B) automatically adjusts the provisioned read and write capacity based on actual traffic patterns, preventing throttling during peak hours without manual intervention. This is the correct action because it dynamically increases capacity when demand spikes and reduces it during low traffic, directly addressing the throttling issue.

Exam trap

The trap here is that candidates often confuse increasing provisioned capacity (Option C) as the only solution, but the exam tests whether you understand that auto scaling (Option B) is the correct managed approach, and that DAX (Option D) can reduce read throttling by caching, making both B and D valid together.

284
MCQmedium

A media company stores millions of thumbnail images in an Amazon S3 bucket. Analysts run ad hoc queries against the image metadata, which is kept as JSON objects in the same bucket. Query latency is unpredictable and costs are rising because Athena scans large volumes of JSON for every query. The team wants faster queries and lower scan cost while keeping the data in S3 and queryable with SQL. Which change should the data engineer make?

A.Use AWS Glue to crawl the metadata, convert it to Apache Parquet partitioned by date, and register the table in the Data Catalog for Athena queries.
B.Enable S3 Intelligent-Tiering on the bucket to automatically move infrequently accessed metadata objects to cheaper storage.
C.Move the metadata into an Amazon DynamoDB table and have analysts query it with PartiQL.
D.Increase the Athena workgroup data usage control limit so queries can scan more data without failing.
AnswerA

Converting JSON metadata to columnar Parquet lets Athena read only the referenced columns and benefits from compression, while date partitioning restricts each query to relevant folders. Registering the table in the Data Catalog makes it directly queryable, reducing scanned bytes and cost while keeping the data in S3.

Why this answer

The root cause is that Athena must read entire JSON objects and every partition for each query. Converting metadata to columnar Parquet and partitioning by date lets the engine read only needed columns and folders, cutting scanned bytes and latency. Catalog registration keeps the data in S3 and preserves SQL access, which matches all stated constraints.

Exam trap

The trap here is assuming that cheaper S3 storage tiers or higher query limits reduce Athena scan cost, when the real lever is the data format and partitioning.

285
MCQmedium

A data engineer is deploying an Amazon Redshift cluster that must be accessible only from within a private VPC and must not have a public IP address. The cluster will be queried by an Amazon EMR cluster in the same VPC and by on-premises BI tools over a VPN connection. Which configuration should the engineer choose?

A.Launch the Redshift cluster in a private subnet group and configure an AWS Site-to-Site VPN connection to the VPC.
B.Launch the Redshift cluster with the publicly accessible setting enabled and attach it to a public subnet group.
C.Launch the Redshift cluster with the publicly accessible setting disabled and attach it to a public subnet group.
D.Launch the Redshift cluster with the publicly accessible setting disabled and attach it to a private subnet group.
AnswerD

Disabling the publicly accessible setting prevents the cluster from receiving a public IP address. Placing it in a private subnet group ensures that only resources within the VPC, such as the EMR cluster, and connected networks over VPN can reach it. This meets the requirement for private-only access without exposing the cluster to the internet.

Why this answer

To ensure a Redshift cluster has no public IP and is reachable only privately, the engineer must disable the publicly accessible setting and use a private subnet group. This prevents internet exposure while allowing access from within the VPC and over VPN. The other options either enable public access or do not explicitly disable it, failing the requirement.

Exam trap

The trap here is assuming that placing a cluster in a private subnet automatically disables public accessibility, when the publicly accessible setting must be explicitly turned off to avoid a public IP.

286
MCQmedium

A data engineer is configuring an Amazon S3 bucket that stores sensitive customer records for analytics. The security team requires that all data be encrypted at rest with keys that are rotated automatically every year and that access be auditable per key. The engineer must minimize operational overhead. Which encryption configuration should be used?

A.SSE-S3 with bucket key enabled
B.SSE-C with customer-provided keys stored in AWS Secrets Manager
C.SSE-KMS with a customer managed key with automatic rotation enabled
D.SSE-KMS with an AWS managed key (aws/s3)
AnswerC

A customer managed KMS key supports enabling automatic key rotation on an annual schedule, and every use of the key is recorded in AWS CloudTrail for auditing. Object encryption at rest is enforced through SSE-KMS, and the data engineer retains control of the key policy. This satisfies the rotation, auditability, and low operational overhead requirements without managing raw key material.

Why this answer

The requirement combines automatic annual key rotation, per-key auditability, and minimal operational overhead. A KMS customer managed key with automatic rotation enabled provides all three: S3 encrypts objects server-side, CloudTrail records key usage, and rotation occurs yearly without manual intervention. AWS managed keys rotate on a different schedule and offer less policy control, while SSE-S3 and SSE-C lack the needed audit or rotation behavior.

Exam trap

The trap here is assuming that SSE-S3 or an AWS managed KMS key provides configurable annual rotation and per-key auditing, when only a customer managed KMS key with rotation enabled meets both requirements.

287
MCQhard

A company uses Amazon DynamoDB to store user session data. The table has a partition key of user_id and a sort key of session_start. The workload is read-heavy and eventually consistent reads are acceptable. The table is provisioned with 1000 RCUs and 500 WCUs. During peak hours, the application experiences throttling on read operations, but CloudWatch shows that the consumed read capacity is well below the provisioned amount. What is the most likely cause of the throttling?

A.The table's sort key is causing uneven data distribution across partitions, leading to throttling.
B.The application is using strongly consistent reads, which consume twice the read capacity units and cause throttling.
C.The table has a hot partition because user_id values are not evenly distributed, causing some partitions to exceed their read capacity limits.
D.The table's provisioned read capacity is too low for the workload, and the engineer should increase it to resolve throttling.
AnswerC

DynamoDB partitions have a maximum read capacity of 3000 RCUs per partition. If a few user_id values are accessed much more frequently than others, those items reside on the same partition, causing that partition to throttle even though overall table capacity is underutilized. This is a classic hot partition scenario, and the fix is to distribute the workload more evenly, such as by adding a random suffix to the partition key.

Why this answer

DynamoDB throttling can occur even when overall consumed capacity is below provisioned if a single partition exceeds its per-partition limit. With a partition key like user_id, if some users are much more active, their items concentrate on one partition, causing that partition to throttle. The solution is to distribute the workload more evenly, such as by adding a random suffix to the partition key or using a composite partition key.

Exam trap

The trap here is assuming that throttling always means insufficient provisioned capacity, overlooking per-partition limits that cause hot partitions.

288
MCQeasy

A data engineer needs to store JSON documents that are frequently read and written by a web application. The data has a flexible schema and requires low-latency queries on primary key lookups. Which AWS service is MOST suitable?

A.Amazon Redshift
B.Amazon S3
C.Amazon DynamoDB
D.Amazon RDS for MySQL
AnswerC

Amazon DynamoDB stores JSON as native documents and delivers single-digit-millisecond latency on primary key lookups, satisfying the low-latency requirement. Its schemaless key-value design accommodates the flexible schema, while provisioned throughput sustains the frequent reads and writes the web application generates.

Why this answer

Amazon DynamoDB is the most suitable service because it is a NoSQL key-value and document database that provides single-digit millisecond latency for primary key lookups, supports flexible schemas for JSON documents, and is designed for high-throughput read/write workloads from web applications. Its fully managed nature and auto-scaling capabilities align with the requirement for frequent, low-latency queries on a flexible schema.

Exam trap

The trap here is that candidates may confuse Amazon S3's ability to store JSON documents with the need for low-latency primary key lookups, overlooking that S3 is not a database and lacks the indexing and query performance required for frequent, transactional reads and writes.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a columnar data warehouse optimized for complex analytical queries on structured data, not for low-latency primary key lookups on JSON documents with frequent writes. Option B is wrong because Amazon S3 is an object storage service that does not support low-latency primary key lookups or native querying without additional services like Athena or S3 Select, and it is not designed for frequent, transactional read/write operations. Option D is wrong because Amazon RDS for MySQL is a relational database with a fixed schema, requiring schema changes for flexible JSON documents, and while it can handle JSON, it does not match DynamoDB's single-digit millisecond latency for primary key lookups at scale.

289
MCQmedium

A data engineer needs to store and analyze time-series data from IoT devices. The data volume is 10 GB per day, and the queries are mostly on the most recent 7 days of data. The engineer wants to minimize storage costs while retaining historical data for 1 year. Which combination of AWS services is most cost-effective?

A.Amazon Timestream
B.Amazon DynamoDB with TTL and S3 for archival
C.Amazon Redshift
D.Amazon RDS with MySQL
AnswerA

Timestream is cost-effective for time-series data with automatic storage tiering.

Why this answer

Amazon Timestream is purpose-built for time-series data, offering automatic tiering between in-memory (for recent 7 days) and magnetic stores (for historical data up to 1 year). This matches the query pattern (mostly recent 7 days) and retention requirement (1 year) while minimizing storage costs through its serverless, pay-per-query model. Timestream also supports time-series-specific functions like interpolation and smoothing, making it more efficient than general-purpose databases for this workload.

Exam trap

The trap here is that candidates often choose DynamoDB with TTL and S3 for archival (Option B) because it seems cost-effective, but they overlook the operational complexity and query latency of accessing historical data in S3, which violates the 'minimize storage costs while retaining historical data for 1 year' requirement without considering query patterns.

How to eliminate wrong answers

Option B (DynamoDB with TTL and S3 for archival) is wrong because DynamoDB is optimized for key-value and document workloads, not time-series analytics; TTL only deletes old data, but querying historical data from S3 requires additional services like Athena or Glue, increasing complexity and latency. Option C (Amazon Redshift) is wrong because Redshift is a columnar data warehouse designed for large-scale analytical queries on structured data, but it is over-provisioned and costly for 10 GB/day of time-series data, and its storage and compute are not optimized for time-series-specific operations like downsampling or retention policies. Option D (Amazon RDS with MySQL) is wrong because RDS is a relational database with fixed storage and compute, leading to higher costs for storing 3.65 TB of historical data (10 GB/day × 365 days) and poor query performance on time-series data without built-in time-series features like automatic retention or partitioning.

290
MCQeasy

A company is using Amazon S3 for data lake storage. They need to query the data directly using SQL without loading it into a database. Which AWS service should be used?

A.Amazon Redshift Spectrum
B.Amazon Athena
C.Amazon EMR
D.AWS Glue
AnswerB

Athena queries data in place in Amazon S3 using standard SQL, requiring no loading into a database. This directly satisfies the requirement to query S3 data lake content with SQL while avoiding ETL or database provisioning.

Why this answer

Amazon Athena is the correct choice because it is a serverless, interactive query service that allows you to analyze data directly in Amazon S3 using standard SQL, without needing to load or transform the data into a database. Athena uses Presto under the hood and supports querying structured, semi-structured, and unstructured data formats (e.g., CSV, JSON, Parquet, ORC) stored in S3, making it ideal for ad-hoc SQL queries on a data lake.

Exam trap

The trap here is that candidates often confuse AWS Glue's data cataloging and ETL capabilities with direct SQL querying, or they assume Redshift Spectrum is a standalone service rather than a feature requiring an existing Redshift cluster, leading them to pick a wrong answer that requires additional infrastructure or is not a query engine.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift Spectrum is a feature of Amazon Redshift that allows querying data in S3 from within a Redshift data warehouse, but it requires an existing Redshift cluster and is not a standalone service for directly querying S3 data without a database. Option C is wrong because Amazon EMR is a big data platform that uses frameworks like Apache Spark, Hive, or Presto for querying S3 data, but it requires provisioning and managing clusters, which adds complexity and is not a serverless SQL-only solution. Option D is wrong because AWS Glue is a serverless data integration service primarily used for ETL (extract, transform, load) jobs and data cataloging, not for directly querying S3 data with SQL; while it can prepare data for Athena, it is not a query engine itself.

291
MCQeasy

A company wants to use Amazon Redshift Spectrum to query data in Amazon S3. The data is in Parquet format and partitioned by date. Which step is required to enable Redshift Spectrum?

A.Load the data into Redshift tables using the COPY command.
B.Create an external schema and external table in the AWS Glue Data Catalog.
C.Create a separate Redshift Spectrum cluster.
D.Copy the data from S3 to Redshift-managed storage.
AnswerB

Redshift Spectrum queries S3 data through external tables registered in the AWS Glue Data Catalog, which you reference via an external schema in the Redshift cluster. This is the mandatory configuration step enabling Spectrum to read the partitioned Parquet data.

Why this answer

Redshift Spectrum allows querying data directly in Amazon S3 without loading it into Redshift. To use Spectrum, you must define an external schema and external table in the AWS Glue Data Catalog (or an external Hive metastore) that points to the S3 location and specifies the Parquet format and partition structure. This enables Redshift to read the data in place using the Spectrum engine.

Exam trap

The trap here is that candidates assume Redshift Spectrum requires a separate cluster or that data must be loaded into Redshift, confusing Spectrum with traditional Redshift ingestion methods like COPY or CTAS.

How to eliminate wrong answers

Option A is wrong because the COPY command loads data into Redshift-managed storage, which bypasses Spectrum's external query capability and incurs storage costs; Spectrum queries data directly from S3 without loading. Option C is wrong because Redshift Spectrum does not require a separate cluster; it runs on the existing Redshift cluster's compute nodes, leveraging the Spectrum layer to access S3. Option D is wrong because copying data from S3 to Redshift-managed storage defeats the purpose of Spectrum, which is to query data in place without moving it.

292
MCQeasy

A company needs to store files that are accessed by multiple EC2 instances in a VPC. The files must be concurrently accessible and durable. Which storage solution should the data engineer choose?

A.Amazon EC2 instance store
B.Amazon Simple Storage Service (Amazon S3)
C.Amazon Elastic Block Store (Amazon EBS)
D.Amazon Elastic File System (Amazon EFS)
AnswerD

Amazon EFS provides a shared, elastic NFS file system that many EC2 instances mount concurrently across Availability Zones, delivering the concurrent access and durability the scenario demands. Instance store and EBS volumes attach to a single instance, so they cannot satisfy the multi-instance requirement.

Why this answer

Amazon EFS provides a fully managed, scalable, and elastic NFS file system that can be concurrently accessed by multiple EC2 instances across multiple Availability Zones. It is designed for high durability (11 nines of durability) and automatically replicates data across multiple AZs within a region, meeting the requirements for concurrent access and durability.

Exam trap

The trap here is that candidates often confuse Amazon EBS Multi-Attach with a general-purpose shared file system, but EBS Multi-Attach is limited to specific io1/io2 volumes, requires cluster-aware applications, and does not provide the POSIX file system semantics or cross-AZ durability that EFS offers.

How to eliminate wrong answers

Option A is wrong because EC2 instance store provides ephemeral block storage that is physically attached to the host; it is not durable (data is lost on instance stop/termination) and cannot be shared concurrently across multiple EC2 instances. Option B is wrong because Amazon S3 is an object storage service, not a file system; it does not support standard file-level locking or NFS/SMB protocols required for concurrent file access from multiple EC2 instances without additional gateways or software. Option C is wrong because Amazon EBS provides block-level storage volumes that can only be attached to a single EC2 instance at a time (except for multi-attach EBS io1/io2 volumes, which are limited to specific instance types and have strict constraints, not a general solution for concurrent file access).

293
MCQhard

A data engineer manages an Amazon DynamoDB table for order events. Reads and writes are evenly spread across a partition key with very high cardinality, but during flash sales the table throttles with ProvisionedThroughputExceededException even though consumed capacity is below the provisioned total. Which cause is MOST likely?

A.The table's on-demand capacity mode must be enabled to remove per-partition limits.
B.The table's items are too large, and each write consumes more than one write capacity unit.
C.The table has too many global secondary indexes, which each consume write capacity.
D.A single partition key value is receiving a disproportionate share of traffic, creating a hot partition that exceeds its per-partition throughput limit.
AnswerD

DynamoDB divides a table into partitions, and each partition has a hard ceiling of roughly 3,000 read units and 1,000 write units per second regardless of the table's total provisioned capacity. If one key value, such as a default or shared customer ID, receives a burst of traffic, that partition saturates and throttles even though aggregate consumed capacity sits below the table's provisioned total.

Why this answer

DynamoDB throughput is provisioned at the table level but enforced at the partition level, and each partition can serve only about 1,000 write units per second. When traffic concentrates on one partition key value, that partition throttles while the table still shows unused capacity. Even distribution across a high-cardinality key normally prevents this, so the flash-sale burst must be funneling requests to a single key value.

Exam trap

The trap here is reading ProvisionedThroughputExceededException as proof that the table needs more capacity, when spare table capacity plus throttling actually indicates a hot partition.

294
MCQeasy

A data engineering team is using AWS Glue to catalog data in an S3 data lake. They have a Glue crawler that runs daily to update the Data Catalog. Recently, they noticed that the crawler is taking longer to run and sometimes fails because of a timeout. The team suspects the issue is due to the large number of small files in the S3 bucket. They need to improve crawler performance and reliability. Which solution should they implement?

A.Configure the crawler to use a different classifier.
B.Use AWS Glue ETL to consolidate small files into larger ones before crawling.
C.Increase the crawler timeout to 24 hours.
D.Schedule the crawler to run more frequently to avoid large data accumulation.
AnswerB

Consolidating many small S3 objects into fewer larger files via Glue ETL reduces the metadata and listing overhead the crawler processes each run. This directly addresses the small-file volume causing slow runs and timeouts, restoring crawler performance and reliability.

Why this answer

Consolidating small files into larger ones (e.g., using AWS Glue ETL with a groupFiles or groupSize option, or a separate compaction job) reduces the number of objects the crawler must list and sample. This directly addresses the root cause: a high volume of small files increases metadata operations and can cause crawler timeouts. By reducing file count, the crawler can complete within the default 24-hour timeout and avoid failures.

Exam trap

The trap here is that candidates assume increasing the timeout or running the crawler more frequently will fix performance issues, but the real bottleneck is the sheer number of small files, which requires data compaction to resolve.

How to eliminate wrong answers

Option A is wrong because changing the classifier affects how the crawler interprets data format (e.g., JSON vs. Parquet), not the number of files or the performance bottleneck caused by small files. Option C is wrong because increasing the timeout to 24 hours does not solve the underlying issue of excessive small files; the crawler may still fail due to resource limits or S3 request throttling, and the default timeout is already 24 hours.

Option D is wrong because running the crawler more frequently would only accumulate more small files over time, worsening the problem and increasing the likelihood of timeouts.

295
Multi-Selecteasy

A data engineer is setting up Amazon S3 bucket policies for a data lake. Which TWO statements are true regarding S3 bucket policies? (Choose TWO.)

Select 2 answers
A.Bucket policies can grant access to accounts in other AWS Organizations
B.Bucket policies are the only way to control access to S3
C.Bucket policies can be applied to individual objects
D.The Principal element in a bucket policy is optional
E.Bucket policies are written in JSON format
AnswersA, E

Cross-account access can be granted via bucket policies.

Why this answer

S3 bucket policies can grant cross-account access to principals in other AWS accounts, including those in different AWS Organizations, by specifying the target account ID or organization ID in the Principal element. This enables centralized data lake access management across organizational boundaries without requiring IAM roles or resource-based policies in each account.

Exam trap

The trap here is that candidates often confuse bucket policies with IAM policies, mistakenly thinking the Principal element is optional in bucket policies (it is required), or that bucket policies can target individual objects (they cannot; they use prefix or tag conditions instead).

296
MCQeasy

A company uses Amazon S3 as its data lake. A data engineer needs to enforce encryption of data at rest using server-side encryption with AWS KMS. Which S3 bucket property should be configured?

A.Default encryption
B.Server access logging
C.Versioning
D.Bucket policy
AnswerA

Configuring default encryption on the bucket applies SSE-KMS automatically to every object written, without relying on request headers. This enforces encryption at rest using AWS KMS as the stem demands, covering uploads from any client or SDK.

Why this answer

Configuring default encryption on an S3 bucket ensures that all objects stored in the bucket are encrypted at rest using server-side encryption. When AWS KMS is specified as the encryption type, S3 automatically encrypts objects with a KMS key (SSE-KMS) upon upload, even if the upload request does not include encryption headers. This enforces encryption at rest without requiring changes to client applications.

Exam trap

The trap here is that candidates often confuse bucket policies (which can enforce encryption conditions) with default encryption (which actually applies encryption), leading them to select bucket policy as the answer when the question asks for the property that enforces encryption of data at rest.

How to eliminate wrong answers

Option B is wrong because server access logging records requests made to the bucket for auditing purposes, but it does not enforce or configure encryption of data at rest. Option C is wrong because versioning preserves, retrieves, and restores every version of every object in the bucket, but it has no effect on encryption settings. Option D is wrong because a bucket policy can deny unencrypted uploads using a condition key like `s3:x-amz-server-side-encryption`, but it does not itself configure the encryption mechanism; it only enforces a policy requirement, whereas default encryption directly applies encryption to all objects.

297
MCQeasy

Refer to the exhibit. A data engineer creates an Amazon Redshift table with the above DDL. The engineer runs a query to find all orders for a specific customer within a date range. Which statement about query performance is correct?

A.The query will be inefficient because the distribution key is not the same as the sort key.
B.The table should use DISTSTYLE EVEN to improve performance.
C.The query will benefit from both the distribution key and the sort key to minimize data scanned.
D.The sort key will not help because the query filters on customer_id first.
AnswerC

Distribution reduces data movement, sort key reduces data scanned.

Why this answer

The DDL defines customer_id as the distribution key and order_date as the sort key. When the query filters on both customer_id (distribution key) and order_date (sort key), Redshift can use partition pruning via the sort key to skip blocks that don't match the date range, and the distribution key ensures that data for the same customer is co-located on the same node slice, minimizing data movement. This combination reduces the amount of data scanned and improves query performance.

Exam trap

The trap here is that candidates assume the sort key is useless if the filter does not start with the sort key column, but Redshift's zone map pruning works on any column in the sort key, and the distribution key filter can still leverage co-location to reduce data movement.

How to eliminate wrong answers

Option A is wrong because the distribution key and sort key do not need to be the same; they serve different purposes—distribution key optimizes data locality for joins and aggregations, while sort key optimizes range-restricted scans. Option B is wrong because DISTSTYLE EVEN distributes rows randomly across slices, which would scatter a single customer's data across all nodes, increasing network traffic and reducing the benefit of the sort key for range scans. Option D is wrong because the sort key on order_date still helps even though the query filters on customer_id first; Redshift can apply predicate-based block pruning on the sort key after the distribution key filter narrows the relevant slices, and the sort key order (customer_id, order_date) means the date filter can still be used efficiently within each customer's data.

298
Multi-Selectmedium

A data engineer is optimizing an Amazon RDS for MySQL database that experiences high write throughput. The engineer wants to improve write performance and reduce latency. Which TWO database-level configuration changes can help achieve this?

Select 2 answers
A.Use Provisioned IOPS (io1 or io2) storage.
B.Reduce the backup retention period to 1 day.
C.Increase the DB instance class to a larger size.
D.Create a Read Replica to offload writes.
E.Enable Multi-AZ for high availability.
AnswersA, C

Provisioned IOPS provides consistent low-latency writes.

Why this answer

Provisioned IOPS (io1 or io2) storage delivers consistent and predictable I/O performance by guaranteeing a specified number of I/O operations per second, which directly reduces latency and improves write throughput for high-write workloads. This is the most effective storage-level change for write-intensive RDS for MySQL databases.

Exam trap

The trap here is that candidates often confuse Multi-AZ with performance improvement, but Multi-AZ is designed for durability and failover, not for speeding up writes.

299
MCQhard

A company has an Amazon DynamoDB table with a provisioned write capacity of 1000 WCU. During a flash sale, the write traffic spikes to 5000 WCU for 10 minutes. The table is not auto-scaled. Which action should the data engineer take to handle the spike without throttling?

A.Convert the table to on-demand capacity mode before the sale.
B.Set a CloudWatch alarm to increase provisioned capacity when write throttling occurs.
C.Use DynamoDB Accelerator (DAX) to cache writes.
D.Enable auto-scaling with a target utilization of 70% and a maximum capacity of 5000 WCU.
AnswerA

On-demand mode removes provisioned WCU entirely, so DynamoDB instantly accommodates the 5000 WCU burst without throttling. Since the table is not auto-scaled and 1000 WCU is fixed, switching before the flash sale is the only action that absorbs the spike.

Why this answer

The table is currently provisioned with 1000 WCU and cannot handle a spike to 5000 WCU. Converting to on-demand mode before the sale allows DynamoDB to automatically handle varying traffic without throttling, as on-demand capacity scales instantly to meet demand. Option D (auto-scaling) might not react quickly enough for a short 10-minute spike, and the table is not currently auto-scaled.

Option C is incorrect because DAX is a read cache and does not buffer or improve write capacity. Option B is reactive and would not prevent initial throttling.

Exam trap

Candidates often assume DAX can handle write spikes because it is a cache, but DAX only caches reads and does not buffer writes. The correct approach is to use on-demand capacity for unpredictable traffic spikes.

How to eliminate wrong answers

Option A is wrong because converting to on-demand capacity mode before the sale would handle the spike without throttling, as on-demand scales instantly to any traffic, but the question's answer key incorrectly marks C as correct. Option B is wrong because setting a CloudWatch alarm to increase provisioned capacity when write throttling occurs is reactive and will cause throttling before the alarm triggers and capacity increases. Option D is wrong because enabling auto-scaling with a target utilization of 70% and a maximum capacity of 5000 WCU would work if configured in advance, but the table is not auto-scaled and the spike is sudden; auto-scaling has a cooldown period and cannot react instantly to a 10-minute spike.

300
Multi-Selectmedium

A data engineer is migrating a large Oracle data warehouse to Amazon Redshift. The engineer needs to ensure optimal performance. Which TWO practices should the engineer follow?

Select 2 answers
A.Choose appropriate sort keys based on common query patterns.
B.Design the schema as a normalized star schema with row-based storage.
C.Manually define compression encodings for each column.
D.Stage data in Amazon S3 before loading into Redshift.
E.Use DISTKEY to distribute data evenly across nodes.
AnswersA, E

Sort keys reduce the amount of data scanned.

Why this answer

Amazon Redshift uses sort keys to physically order data on disk, which allows the query optimizer to skip large blocks of data during scans via zone maps. Choosing sort keys based on common query patterns (e.g., range filters or frequent GROUP BY columns) dramatically reduces I/O and improves query performance, especially for large tables.

Exam trap

The trap here is that candidates often confuse Redshift's columnar storage with row-based storage and assume a normalized star schema is optimal, when in fact Redshift is designed for denormalized, columnar tables with explicit sort and distribution keys.

← PreviousPage 4 of 5 · 358 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Data Store Management questions.