Build a correct pipeline: filter early with $match on indexed fields, group with the right accumulator like $avg, and use $unionWith to append another collection's results. The single most important thing is stage order—$match and $sort belong before $group to stay efficient.
Start practicing
Aggregation Framework — choose a session length
Free · No account required
Domain overview
The Aggregation Framework domain covers building multi-stage pipelines that transform, filter, group, and combine documents in MongoDB. Questions present realistic scenarios—session analytics, indexed filtering, cross-collection merging—and ask you to choose the correct stage, operator, or ordering. Expect a mix of operator identification and pipeline-design reasoning rather than syntax memorization.
Exam objectives
Selecting stages like $match, $group, $sort, $project, and $unionWith for a given transformation goal
Using accumulators such as $avg, $sum, $min, and $max inside a $group stage
Ordering pipeline stages so $match and $sort leverage indexes before expensive operations
Combining collections with $unionWith and shaping output with $project and $addFields
Placing $match after $group, which forces processing of unfiltered documents and prevents early index use.
Confusing $avg with $sum or applying accumulators outside a $group stage where they are invalid.
Assuming $unionWith deduplicates or joins on a key; it appends result sets and does not match documents.
Click any question to see the full explanation and answer options, or start a focused practice session above.
Which stage should be placed first in an aggregation pipeline to optimize performance when filtering a large collection based on an indexed field?
2Which TWO stages are considered 'blocking' stages that can potentially consume excessive memory if not managed correctly?
3Which operator is used within a $project stage to calculate the square root of a numerical field?
4In an aggregation pipeline, what is the purpose of the $addFields stage?
5What happens if an aggregation stage exceeds the 100MB memory limit and 'allowDiskUse' is set to false?
6Which operator is used to calculate the average value of a numeric field in a $group stage?
7Which operator is used to perform a left outer join between two collections in an aggregation pipeline?
8Which operator would you use to filter for documents where an array field contains ALL of the elements in a specified array?
9What is the primary benefit of using $unionWith to combine data from two collections in a single pipeline?
10Which aggregation stage allows you to replace the root of the document with a specified document or field?
11Refer to the exhibit. Which stage should be used to transform this document into three separate documents, one for each tag?
12Why does using a $limit stage immediately after a $sort stage provide a performance advantage?
13Which operator is most appropriate to concatenate an array of strings into a single string field?
14Refer to the exhibit. Why is this pipeline considered suboptimal?
15What is the purpose of the $accumulator stage in the aggregation framework?
16What happens if a $sort stage exceeds the 100MB RAM limit without an index?
17You are processing a large collection of user sessions. You need to calculate the average duration of sessions grouped by 'userId', but you must filter out any sessions lasting less than 5 seconds before performing the grouping. Which stage order is most efficient?
18Which aggregation operator is used to transform the shape of documents by adding, renaming, or excluding fields?
19Which TWO of the following are valid accumulators that can be used within a $group stage?
20Refer to the exhibit. What is the effect of the $unwind stage on the aggregation pipeline?
21When using the $lookup stage, what happens if the 'foreignField' contains values that are not present in the 'localField' of the joined collection?
22Which stage would you use to limit the output of your aggregation to the top 10 most expensive items?
23What is the purpose of the $merge stage in an aggregation pipeline?
24Which aggregation operator is most appropriate for counting the total number of documents in a collection?
25Refer to the exhibit. What happens to documents in the collection that do not contain the 'category' field?
26What is the primary benefit of using an index with the $match stage at the beginning of an aggregation pipeline?
27Which stage is used to change the order of documents in an aggregation pipeline?
28Refer to the exhibit. What is the expected behavior if the 'price' field exists but the 'tax' field is missing in a document?
29An application requires grouping orders by customer ID and calculating the total purchase amount per customer. However, some customers have orders with missing or null purchase amounts that must be excluded before calculation. Which aggregation pipeline stage sequence correctly achieves this goal?
30An analytics collection contains order documents with an array of items. You need to output a single document for each element in the items array while preserving documents that have an empty or missing items array. Which stage configuration achieves this?
31A developer is building an aggregation pipeline on a collection of sensor readings. The pipeline must group documents by sensorId and compute two values: the total number of readings and the timestamp of the most recent reading. Which $group accumulator pair should be used?
32A retail company stores orders in a collection where each document has a `customerId` field and a `total` field. The analytics team needs to see, for each customer, both the sum of all order totals and a list of the individual order totals. Which aggregation pipeline stage and accumulator combination should be used?
33You maintain a MongoDB collection named `readings` that stores hourly sensor measurements. Each document includes a `sensorId` string, a `recordedAt` date, and a `temperature` numeric value. You need to produce a report that, for each sensor, lists the sensor identifier alongside the timestamp of its single highest temperature reading. Duplicate temperatures are possible, and in that case any one of the tied readings is acceptable. Which aggregation pipeline should you run?
34A developer needs to calculate the total revenue per product category from a sales collection. The collection has documents with fields `category` and `amount`. Which aggregation pipeline stage should be used to compute the sum of amounts for each category?
35A support team stores tickets in a `tickets` collection where each document has a `status` field that may be missing on legacy records. A developer runs a pipeline whose first stage is { $match: { status: { $ne: "closed" } } } and expects every non-closed ticket to appear in the results. Which set of documents does this stage pass downstream?
36An aggregation pipeline processes a collection of sensor readings with fields `sensorId`, `timestamp`, and `value`. The pipeline must output, for each sensor, the most recent reading's value and the average value across all readings. Which pipeline correctly achieves this?
37You have a collection of sensor readings where each document includes a timestamp field `recordedAt` and a temperature field `tempC`. You need to generate a report that groups readings into 15-minute intervals and calculates the average temperature per interval. Which aggregation expression should you use to create the grouping key?
38You are building an aggregation pipeline that processes millions of documents. The pipeline includes a $match stage, a $group stage that calculates sums, and a $sort stage. You notice the pipeline is spilling to disk and running slowly. Which stage should you place first to optimize performance while maintaining correct results?
39A developer is building an aggregation pipeline on a collection of blog posts. Each post document contains an array field `comments`, where each comment is an object with `author` and `text`. The developer needs to produce a separate document for each comment, preserving the post's `title` and the comment's `author` and `text`. Which TWO stages are required to achieve this? (Choose two.)
40An analyst needs a pipeline that reads from a `sales` collection and, for each document, emits one output document per entry in the `regions` array, copying all other fields unchanged. Which stage accomplishes this?
41You are troubleshooting a pipeline over a large `telemetry` collection that exceeds the aggregation memory limit while running on a single mongod instance. You must let the pipeline complete successfully without changing the server-wide configuration. Which two actions will allow the pipeline to keep running? (Choose two.)
42You are building an aggregation pipeline on a collection of sensor readings where each document has a `timestamp` field and a `value` field. You need to compute, for each calendar day, the highest `value` and the average `value` across all readings for that day. Which `$group` stage accomplishes this?
43You are reviewing an aggregation pipeline on an `orders` collection that must return only orders with a `total` greater than 100 and then compute the number of orders per `customerId`. The collection has an index on `total`. Which two stages should you use, and in what order, to maximize performance while returning the correct results? (Choose two.)
44A retail company stores order documents in a collection named orders. Each document has a field `total` (numeric) and a field `status` (string). The analytics team needs to produce a report that shows, for each `status`, the total revenue and the number of orders, but only for orders where `total` is greater than 100. The report must be sorted by total revenue in descending order. Which aggregation pipeline correctly produces this report?
45You are aggregating a `logs` collection where each document has a `timestamp` and a `level` field. You need to produce a single document that contains two arrays: one with all `timestamp` values for `level: "ERROR"` and one with all `timestamp` values for `level: "WARN"`. Which pipeline should you use?
46A financial application stores transactions in a collection named transactions. Each document has fields: `accountId` (string), `amount` (decimal), `type` (string, either 'debit' or 'credit'), and `timestamp` (date). The team needs to calculate the net balance per account, defined as the sum of credits minus the sum of debits. The result should include only accounts with at least 10 transactions. Which aggregation pipeline achieves this?
47A pipeline on a `sales` collection must return only the `item` and `price` fields for all documents where `price` is greater than 50, and it should not return the `_id` field. Which pipeline achieves this?
48A social media platform stores posts in a collection named posts. Each post document contains an array field `comments`, where each comment is a subdocument with fields `user` and `text`. The development team needs to generate a report that lists each comment as a separate document, including the post's `_id` and the comment's `user` and `text`. Which aggregation stage should be used to achieve this?
49A logistics company stores shipment documents in a collection named shipments. Each document has fields: `origin` (string), `destination` (string), `weight` (number), and `status` (string). The analytics team wants to find the total weight of shipments for each unique pair of origin and destination, but only for shipments with status 'delivered'. They also want to include the average weight per pair and sort the results by total weight descending. Which two stages are essential to achieve the grouping and calculation of these metrics? (Choose two.)
Build a correct pipeline: filter early with $match on indexed fields, group with the right accumulator like $avg, and use $unionWith to append another collection's results. The single most important thing is stage order—$match and $sort belong before $group to stay efficient.
The Courseiva C100DEV question bank contains 49 questions in the Aggregation Framework domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Aggregation Framework domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included