A data team needs to extract data from S3 and RDS, transform it (clean, enrich, join), and load it into Amazon Redshift for analytics. They want a serverless service that discovers and catalogues data schemas automatically and runs the ETL jobs without provisioning servers. Which AWS service provides this?
AWS Glue is a serverless ETL service that combines schema discovery, a managed data catalog, and Spark-based transformation jobs in one offering. A Glue Crawler automatically scans S3 or databases, infers schemas, and writes table metadata to the Glue Data Catalog, which makes the data immediately queryable by services like Athena and Redshift Spectrum. Glue ETL Jobs run on a managed, auto-scaling Spark environment without any infrastructure provisioning, and can load transformed results directly into Amazon Redshift, exactly matching the serverless and schema-discovery requirements of the scenario.
Why this answer
AWS Glue is a fully managed, serverless ETL service that automatically discovers and catalogs data schemas using its Crawler feature, which populates the AWS Glue Data Catalog. It can extract data from S3 and RDS, transform it (clean, enrich, join), and load it into Amazon Redshift without any server provisioning or management.
Exam trap
The trap here is that candidates confuse AWS Glue with Amazon EMR because both can run Spark-based ETL, but EMR requires server provisioning and lacks automatic schema discovery, while Glue is fully serverless and includes the Data Catalog.
How to eliminate wrong answers
Option A is wrong because Amazon EMR is a cluster-based big data platform that requires provisioning and managing EC2 instances (servers), and it does not automatically discover or catalog data schemas. Option B is wrong because AWS Data Pipeline is a managed orchestration service but it is not serverless—it relies on EC2 instances or task runners that must be provisioned, and it lacks built-in schema discovery and cataloging. Option D is wrong because Amazon Kinesis Data Firehose is a serverless streaming data ingestion service that loads data into destinations like S3 or Redshift, but it does not perform complex transformations (e.g., joins, enrichment) and has no schema discovery or cataloging capabilities.