Be able to pick the right AWS service for each pipeline stage and justify it on cost, latency, and operational overhead. The single most important thing: know how Athena, Glue, Kinesis, and Redshift are priced and where partitioning, compression, and bookmarks cut cost.
Start practicing
Data Engineering — choose a session length
Free · No account required
Domain overview
The Data Engineering domain covers ingesting, transforming, and storing data for ML workloads on AWS. Questions test service selection and cost/latency trade-offs: Athena over S3, AWS Glue ETL and crawlers, Kinesis and MSK streaming, Amazon Redshift, and S3 storage classes, plus how these feed SageMaker training and inference.
Exam objectives
Choosing Athena, Glue, or EMR for querying and transforming S3 data at scale
Configuring AWS Glue crawlers, Data Catalog tables, and job bookmarks for ETL
Selecting Kinesis Data Streams, Data Firehose, or MSK for streaming ingestion
Using S3 storage classes, partitioning, and compression to reduce scan and storage cost
Assuming Athena charges by data returned rather than data scanned, so unpartitioned or uncompressed queries cost far more than expected.
Forgetting Glue job bookmarks, causing repeated reprocessing of already-transformed S3 objects on each scheduled run.
Picking Kinesis Data Firehose when sub-second, custom-consumer latency is required, since Firehose buffers before delivery.
Practice questions for the Data Engineering domain are being added. Check back soon.
← Back to all MLS-C01 domainsBe able to pick the right AWS service for each pipeline stage and justify it on cost, latency, and operational overhead. The single most important thing: know how Athena, Glue, Kinesis, and Redshift are priced and where partitioning, compression, and bookmarks cut cost.
The Courseiva MLS-C01 question bank contains 0 questions in the Data Engineering domain, covering the 20% of the exam attributed to this domain in the official Amazon Web Services blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data Engineering domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included