A company is designing a data lake on Amazon S3. The data includes personal identifiable information (PII). The data engineer must ensure that only authorized users can access the data, and that access is logged for auditing. Which combination of services should the data engineer use?
S3 bucket policies and IAM policies together enforce least-privilege authorisation on the PII objects, while CloudTrail data events capture object-level API activity such as GetObject, satisfying the auditing requirement that management events alone would not record.
Why this answer
S3 bucket policies combined with IAM policies provide fine-grained access control to restrict who can access the data, while AWS CloudTrail with data events logs every S3 object-level operation (e.g., GetObject, PutObject) for auditing. This combination directly meets the requirements of authorized access and logging for PII data.
Exam trap
The trap here is that candidates often assume CloudTrail automatically logs all S3 operations, but it only logs management events by default; data events must be explicitly enabled, and many overlook this distinction when designing for auditing.
How to eliminate wrong answers
Option B is wrong because Amazon S3 access points and VPC endpoints control network-level access and simplify bucket management, but they do not provide the required logging of data access for auditing. Option C is wrong because Amazon Macie discovers and classifies PII, and S3 Object Lock prevents deletion or overwriting, but neither service enforces access control or logs data access events. Option D is wrong because AWS KMS encrypts data at rest, which protects confidentiality but does not control who can access the data, and while AWS CloudTrail logs API calls, it does not log data events by default; without enabling data event logging, object-level access (e.g., reading a file) is not recorded.