A solutions architect is designing a high-performance computing (HPC) workload on AWS that requires a shared POSIX-compliant file system with high throughput and low latency for thousands of concurrent compute instances. The workload is temporary, running for a few hours each week, and the team wants to minimize cost. Which storage solution should the architect recommend?
FSx for Lustre is designed for HPC workloads, providing a POSIX-compliant file system with sub-millisecond latencies and hundreds of GB/s throughput. A scratch file system is cost-effective for temporary workloads and does not replicate data, which suits the few-hours-per-week pattern. Linking to S3 allows seamless data import and export, reducing manual data movement.
Why this answer
FSx for Lustre is purpose-built for HPC, offering a POSIX-compliant, high-throughput, low-latency file system. A scratch deployment is cost-effective for temporary workloads and can be linked to S3 for data staging. EFS, S3 with FUSE, and FSx for Windows File Server do not provide the required performance or protocol compatibility for this scenario.
Exam trap
The trap here is assuming that Amazon EFS with Provisioned Throughput can match the performance of FSx for Lustre for HPC, when Lustre is specifically optimized for high-throughput, low-latency parallel file access.