Courseiva

CCNA Aggregation Framework Questions

49 questions · Aggregation Framework · All types, answers revealed

1
MCQmedium

A social media platform stores posts in a collection named posts. Each post document contains an array field `comments`, where each comment is a subdocument with fields `user` and `text`. The development team needs to generate a report that lists each comment as a separate document, including the post's `_id` and the comment's `user` and `text`. Which aggregation stage should be used to achieve this?

A.$unwind
B.$project
C.$replaceRoot
D.$group
AnswerA

The $unwind stage deconstructs an array field from the input documents to output a document for each element. In this scenario, using $unwind on the `comments` array will produce one document per comment, preserving the parent post's fields such as `_id`. This exactly matches the requirement to list each comment as a separate document. Additional stages like $project can then reshape the output to include only the desired fields.

Why this answer

The $unwind stage is specifically designed to deconstruct an array field, outputting one document per array element. In this case, unwinding the `comments` array will produce a document for each comment, retaining the parent post's `_id`. This is the correct stage to flatten the comments array and meet the reporting requirement.

Exam trap

The trap here is confusing $project with $unwind; $project can include the array but does not create separate documents for each element.

2
MCQeasy

Which operator is used to perform a left outer join between two collections in an aggregation pipeline?

A.$join
B.$merge
C.$lookup
D.$unionWith
AnswerC

$lookup is the native aggregation operator that performs a left outer join. It allows you to specify the 'from' collection, the 'localField', the 'foreignField', and the 'as' alias for the resulting array, making it the correct and standard choice for cross-collection data aggregation.

Why this answer

The $lookup operator allows for performing an equality match between a field from the input documents and a field from the documents in another collection. It is the primary method for joining data across collections in MongoDB. Mastering $lookup is essential for developers transitioning from relational databases, as it enables the retrieval of related data while keeping the power of the aggregation framework's transformations.

Exam trap

Students often confuse $lookup with $match or $project, attempting to use projection stages to join data across separate collections.

3
MCQmedium

When using the $lookup stage, what happens if the 'foreignField' contains values that are not present in the 'localField' of the joined collection?

A.The entire document is dropped from the pipeline result.
B.The database throws an error and the aggregation fails.
C.The result field contains an empty array for those documents.
D.The result field is set to null.
AnswerC

$lookup functions as a left outer join. When the foreign side of the join has no matching documents for a specific input, the field specified in the 'as' parameter is populated with an empty array. This allows developers to handle 'missing' data programmatically without losing the primary record.

Why this answer

The $lookup operator acts as a left outer join. If there is no match in the foreign collection for a given document, the resulting array field will simply be empty. This behavior is standard for left outer joins, ensuring that no source documents are dropped from the result set simply because they lack a corresponding record in the joined collection, preserving the integrity of the primary dataset.

Exam trap

Candidates often assume that an unmatched $lookup will drop the parent document entirely, confusing it with an inner join rather than recognizing its left outer join behavior.

4
MCQmedium

What happens if an aggregation stage exceeds the 100MB memory limit and 'allowDiskUse' is set to false?

A.The pipeline execution stalls until memory is freed.
B.The database automatically increases the limit.
C.The aggregation operation fails and returns an error.
D.The pipeline skips the stage and continues.
AnswerC

When memory usage exceeds 100MB and allowDiskUse is false, the database throws an error to prevent resource exhaustion. This forces developers to design more efficient pipelines or intentionally enable disk usage, ensuring that resource-heavy queries are executed with full awareness of the potential performance trade-offs.

Why this answer

By default, MongoDB limits each aggregation stage to 100MB of RAM to prevent a single query from consuming excessive system resources and affecting database stability. If a stage exceeds this limit without explicit permission to write to temporary files on disk, the query will fail with an error. Developers must proactively optimize pipelines to minimize memory footprint or enable disk use when large-scale processing is required.

Exam trap

Candidates often assume that MongoDB will automatically spill to disk when memory is exceeded, forgetting that allowDiskUse must be explicitly set to true to enable this behavior.

5
MCQeasy

Which aggregation operator is most appropriate for counting the total number of documents in a collection?

A.$group with $sum
B.$count
C.$limit
D.$match
AnswerB

$count is the designated stage for returning a count of documents. It simplifies the pipeline by replacing the need for a $group operation, making it the most efficient and readable choice for counting records after applying any necessary filters in the preceding stages of your aggregation.

Why this answer

The $count stage is a specialized aggregation stage that returns a document containing the number of input documents at that point in the pipeline. It is shorthand for a $group stage with a $sum accumulator. Using $count is more readable and concise than building a manual $group stage, making it the best practice for simple document counting tasks within an aggregation pipeline.

Exam trap

Candidates often overcomplicate the task by writing a full $group stage with a $sum accumulator, missing the fact that $count provides a more concise and readable built-in alternative.

6
MCQmedium

Which TWO of the following are valid accumulators that can be used within a $group stage?

A.$sum
B.$match
C.$avg
D.$project
E.$sort
AnswerA, C

$sum is a standard accumulator that calculates the sum of all numerical values in the group. It is essential for tasks like calculating total revenue, total item counts, or tallying user engagement metrics across various categories or time periods within a MongoDB aggregation pipeline.

Why this answer

The $group stage relies on accumulator expressions to perform calculations across a group of documents. Understanding which operators are available is vital for creating summary reports. $sum and $avg are standard accumulators that allow developers to perform arithmetic operations on grouped data, enabling complex analytical queries that reveal trends and metrics directly from the database layer without requiring application-side processing.

Exam trap

Candidates often select non-accumulator operators like $match or $sort, failing to distinguish between pipeline stages and the specific functions permitted within the $group stage.

7
MCQeasy

A pipeline on a `sales` collection must return only the `item` and `price` fields for all documents where `price` is greater than 50, and it should not return the `_id` field. Which pipeline achieves this?

A.[ { $project: { _id: 0, item: 1, price: 1 } }, { $match: { price: { $gt: 50 } } } ]
B.[ { $match: { price: { $gt: 50 } } }, { $project: { item: 1, price: 1 } } ]
C.[ { $match: { price: { $gt: 50 } } }, { $project: { _id: 0, item: 1, price: 1 } } ]
D.[ { $match: { price: { $gt: 50 } } }, { $unset: ["_id", "department"] } ]
AnswerC

The `$match` stage filters documents where `price` is greater than 50, and the `$project` stage includes only `item` and `price` while explicitly excluding `_id` with `_id: 0`. This returns exactly the required fields for qualifying documents. The order is efficient because filtering happens before projection.

Why this answer

The pipeline should filter with `$match` on `price` first, then use `$project` to include `item` and `price` while excluding `_id`. This returns only the requested fields for documents meeting the condition. Projecting before matching or using `$unset` does not satisfy the requirement to return only those two fields without `_id`.

Exam trap

The trap here is forgetting that `_id` is included by default in a `$project` inclusion projection, so it must be explicitly excluded when not wanted.

8
MCQmedium

An application requires grouping orders by customer ID and calculating the total purchase amount per customer. However, some customers have orders with missing or null purchase amounts that must be excluded before calculation. Which aggregation pipeline stage sequence correctly achieves this goal?

A.{ $group: { _id: "$customerId", total: { $sum: "$purchaseAmount" } } }, { $match: { purchaseAmount: { $ne: null } } }
B.{ $project: { customerId: 1, purchaseAmount: 1 } }, { $group: { _id: "$customerId", total: { $sum: "$purchaseAmount" } } }
C.{ $match: { purchaseAmount: { $ne: null } } }, { $group: { _id: "$customerId", total: { $sum: "$purchaseAmount" } } }
D.{ $sort: { purchaseAmount: 1 } }, { $group: { _id: "$customerId", total: { $sum: "$purchaseAmount" } } }
AnswerC

Filtering with `$match` before `$group` removes documents whose `purchaseAmount` is null, satisfying the exclusion constraint, so `$sum` only accumulates valid numeric values per `customerId`. Running `$group` first would still sum nulls as zero and inflate nothing, but the stem explicitly requires exclusion prior to calculation.

Why this answer

Filtering out documents with null or missing purchase amounts using the $match stage ensures that subsequent grouping operations do not process corrupted or incomplete data. Placing $match before $group also allows MongoDB to utilize indexes on the purchaseAmount field, significantly improving aggregation performance on large collections.

Exam trap

Candidates frequently place the filtering condition inside the $group stage using conditional expressions, forgetting that $match early in the pipeline provides critical index optimization benefits.

9
MCQeasy

Which stage is used to change the order of documents in an aggregation pipeline?

A.$order
B.$sort
C.$match
D.$project
AnswerB

$sort is the correct stage used to order documents. It accepts an object defining the fields to sort by and the direction (1 for ascending, -1 for descending). This stage is essential for preparing data for downstream operations that depend on order, such as pagination or ranking.

Why this answer

The $sort stage is the standard operator for ordering the documents in a pipeline based on specified fields. It supports both ascending and descending order. Proper sorting is a requirement for many other stages, such as $limit or $group, when those stages rely on the order of input data to produce correct or expected results in a deterministic manner.

Exam trap

Candidates confuse sorting stages with grouping stages or incorrectly assume documents are automatically returned in a predictable insertion order.

10
MCQmedium

You maintain a MongoDB collection named `readings` that stores hourly sensor measurements. Each document includes a `sensorId` string, a `recordedAt` date, and a `temperature` numeric value. You need to produce a report that, for each sensor, lists the sensor identifier alongside the timestamp of its single highest temperature reading. Duplicate temperatures are possible, and in that case any one of the tied readings is acceptable. Which aggregation pipeline should you run?

A.db.readings.aggregate([ { $group: { _id: "$sensorId", temperatures: { $push: "$temperature" }, times: { $push: "$recordedAt" } } }, { $sort: { temperatures: -1 } } ])
B.db.readings.aggregate([ { $sort: { temperature: -1 } }, { $group: { _id: "$sensorId", recordedAt: { $last: "$recordedAt" }, temperature: { $last: "$temperature" } } } ])
C.db.readings.aggregate([ { $group: { _id: "$sensorId", recordedAt: { $first: "$recordedAt" }, temperature: { $first: "$temperature" } } }, { $sort: { temperature: -1 } } ])
D.db.readings.aggregate([ { $sort: { sensorId: 1, temperature: -1 } }, { $group: { _id: "$sensorId", recordedAt: { $first: "$recordedAt" }, temperature: { $first: "$temperature" } } } ])
AnswerD

The $sort stage orders documents by sensorId ascending and temperature descending, so within each sensor the highest temperature comes first. The subsequent $group with $first then captures that leading document's recordedAt and temperature while grouping by sensorId. This is the canonical top-one-per-group pattern in the aggregation framework and satisfies the requirement exactly.

Why this answer

Selecting the top record per group requires ordering the input before the grouping stage, because accumulators such as $first operate on the incoming document stream order. Sorting by the group key and then by the ranking field descending guarantees the desired record arrives first for each group, letting $first extract both the timestamp and the measurement together.

Exam trap

The trap here is assuming $first or $last selects the minimum or maximum value automatically, when they simply take the first or last document in the current stream order.

11
MCQeasy

Which aggregation operator is used to transform the shape of documents by adding, renaming, or excluding fields?

A.$match
B.$project
C.$group
D.$lookup
AnswerB

$project is explicitly designed to reshape documents. It allows for the inclusion, exclusion, or renaming of fields and the creation of new fields based on expressions. This makes it the standard operator for defining the final structure of the result set emitted from an aggregation pipeline.

Why this answer

The $project stage is the primary tool for shaping documents within an aggregation pipeline. It allows developers to specify which fields should be included in the output, calculate new fields, or reshape existing nested documents. Mastering $project is essential for data sanitization, privacy compliance, and preparing output payloads for application consumption by excluding sensitive internal database fields.

Exam trap

Test-takers frequently confuse document shaping stages like $project with data grouping stages like $group, incorrectly attempting to aggregate total sums or averages using $project.

12
Multi-Selectmedium

Which TWO stages are considered 'blocking' stages that can potentially consume excessive memory if not managed correctly?

Select 2 answers
A.$match
B.$group
C.$sort
D.$project
E.$limit
AnswersB, C

The $group stage must consume all input documents to calculate aggregate results like sums or averages. Because it must see the entire stream to finalize values for grouping keys, it is a blocking operation that requires significant RAM. Developers must be cautious with memory limits when grouping.

Why this answer

Both $group and $sort are blocking stages because they require all preceding documents to be processed before they can produce their first output document. This behavior means they must maintain state in memory or spill to temporary files on disk if the working set is large. Understanding which stages are blocking is essential for developers to avoid pipeline crashes or performance degradation when handling large-scale data aggregations.

Exam trap

Candidates often overlook that $sort and $group are blocking stages, failing to realize that these operations must buffer all incoming documents before they can emit any results downstream.

13
MCQhard

A financial application stores transactions in a collection named transactions. Each document has fields: `accountId` (string), `amount` (decimal), `type` (string, either 'debit' or 'credit'), and `timestamp` (date). The team needs to calculate the net balance per account, defined as the sum of credits minus the sum of debits. The result should include only accounts with at least 10 transactions. Which aggregation pipeline achieves this?

A.db.transactions.aggregate([ { $group: { _id: "$accountId", netBalance: { $sum: { $cond: [ { $eq: [ "$type", "credit" ] }, "$amount", { $multiply: [ "$amount", -1 ] } ] } }, count: { $sum: 1 } } }, { $match: { count: { $gte: 10 } } } ])
B.db.transactions.aggregate([ { $match: { $expr: { $gte: [ { $size: "$transactions" }, 10 ] } } }, { $group: { _id: "$accountId", netBalance: { $sum: { $cond: [ { $eq: [ "$type", "credit" ] }, "$amount", { $multiply: [ "$amount", -1 ] } ] } } } } ])
C.db.transactions.aggregate([ { $group: { _id: "$accountId", netBalance: { $sum: { $cond: [ { $eq: [ "$type", "credit" ] }, "$amount", { $multiply: [ "$amount", -1 ] } ] } }, count: { $sum: 1 } } }, { $match: { count: { $gt: 10 } } } ])
D.db.transactions.aggregate([ { $group: { _id: "$accountId", netBalance: { $sum: { $cond: [ { $eq: [ "$type", "credit" ] }, "$amount", { $multiply: [ "$amount", -1 ] } ] } }, count: { $sum: 1 } } }, { $having: { count: { $gte: 10 } } } ])
AnswerA

This pipeline groups by accountId, using $cond to add amount for credits and subtract amount for debits, effectively computing net balance. It also counts transactions per account. Then it filters groups with count >= 10 using $match. This correctly implements the requirement and is efficient because the grouping is done first, then filtered by count, which is necessary since the count is computed during grouping.

Why this answer

The correct pipeline groups transactions by accountId, computes net balance by conditionally adding credits and subtracting debits, and counts transactions. It then filters groups with a count of at least 10 using $match. This approach correctly handles the arithmetic and the threshold, and it is efficient because grouping is necessary to compute the count per account.

Exam trap

The trap here is using $gt instead of $gte when the requirement says 'at least 10', which includes exactly 10.

14
MCQmedium

Refer to the exhibit. Which stage should be used to transform this document into three separate documents, one for each tag?

A.$group
B.$unwind
C.$project
D.$match
AnswerB

The $unwind stage serves to explode an array into individual documents. By setting preserveNullAndEmptyArrays to false, it generates one document for every item in the array, making it the perfect tool for flattening data structures so that subsequent aggregation stages can process each array element independently.

Why this answer

The $unwind stage is designed specifically to deconstruct an array field from the input documents to output a document for each element. This is essential for operations like counting occurrences of individual array elements or calculating statistics across specific array items. Without $unwind, the $group stage would treat the entire array as a single value rather than individual components, preventing granular analysis of the array contents.

Exam trap

Candidates often confuse $unwind with $group or $project, mistakenly believing that projection or grouping alone can split array elements into separate individual documents.

15
MCQeasy

Which operator is used within a $project stage to calculate the square root of a numerical field?

A.$pow
B.$sqrt
C.$abs
D.$exp
AnswerB

The $sqrt operator is the standard aggregation function for calculating the square root of a non-negative number. It is highly optimized for performance and is the correct choice for mathematical transformations involving square roots within an aggregation pipeline's $project or $addFields stages.

Why this answer

The $sqrt operator is part of the aggregation framework's mathematical expression set. It is specifically designed to compute the square root of a given number or expression. Understanding available mathematical operators allows developers to perform complex calculations directly within the database, minimizing the need to retrieve raw data and process it in the application layer, which improves overall system efficiency.

Exam trap

Candidates often try to use standard JavaScript Math functions like Math.sqrt() inside the aggregation pipeline, forgetting that MongoDB requires specific aggregation operators like $sqrt to perform calculations.

16
MCQmedium

What is the purpose of the $accumulator stage in the aggregation framework?

A.To replace $group.
B.To define custom aggregation logic.
C.To improve performance over native operators.
D.To perform joins across collections.
AnswerB

$accumulator allows for complex, user-defined logic that goes beyond standard operators. By providing an init, accumulate, and merge function, you gain full control over how documents are processed and aggregated, enabling custom calculations that are not supported by standard MongoDB aggregation operators.

Why this answer

The $accumulator stage allows developers to define custom aggregation logic using JavaScript functions. This is useful for scenarios where built-in accumulators like $sum or $avg are insufficient, such as calculating complex statistical models or performing custom data parsing that requires state maintenance during the grouping process. Because it executes JavaScript, it should be used judiciously, as it can be slower than native pipeline operators.

Exam trap

Candidates often assume $accumulator is a native performance-optimized operator, failing to account for the overhead of executing custom JavaScript code, which can significantly slow down large-scale data processing.

17
MCQhard

Refer to the exhibit. What happens to documents in the collection that do not contain the 'category' field?

A.They are excluded from the results.
B.They are assigned to a group with _id: null.
C.The aggregation throws an error.
D.They are grouped individually as unique documents.
AnswerB

When the field specified in the _id expression of a $group stage is missing, MongoDB treats the value as null. All such documents are subsequently aggregated into a single group where the _id field is set to null, effectively capturing all documents that lack the 'category' key.

Why this answer

When grouping by a field, MongoDB treats missing fields or null values as having the value 'null'. Therefore, all documents lacking a 'category' field will be grouped together under the ID 'null'. This behavior is consistent and ensures no documents are silently dropped during a $group operation, allowing developers to identify and handle data quality issues where mandatory fields might be missing from certain records.

Exam trap

Candidates often assume documents missing the grouping field are excluded from the results, not realizing that MongoDB treats missing fields as null and groups them together under the null key.

18
MCQmedium

What is the primary benefit of using an index with the $match stage at the beginning of an aggregation pipeline?

A.It guarantees that the output documents are sorted.
B.It allows the pipeline to bypass the 100MB RAM limit.
C.It enables the engine to quickly select a subset of documents.
D.It automatically formats the output fields.
AnswerC

Indexes allow the database to locate relevant documents quickly, avoiding the need to perform a full collection scan. By using an index in the initial $match stage, the engine processes fewer documents, significantly improving performance and reducing resource consumption for every subsequent stage of the aggregation pipeline.

Why this answer

When an aggregation pipeline starts with a $match stage, MongoDB can utilize indexes to filter the collection before processing any further stages. This drastically reduces the number of documents loaded into memory, which is the most effective way to optimize pipeline performance. Without index usage, MongoDB must perform a collection scan, which is significantly slower as the dataset grows, making index usage essential for production applications.

Exam trap

Test-takers often place $match stages later in the aggregation pipeline, missing the performance benefits of filtering early with indexes.

19
MCQhard

Refer to the exhibit. What is the effect of the $unwind stage on the aggregation pipeline?

A.It removes the 'details' field entirely from the resulting documents.
B.It flattens the 'details' array, creating one output document per array element.
C.It merges all elements of the 'details' array into a single string.
D.It sorts the 'details' array in ascending order.
AnswerB

The $unwind stage processes array fields by creating a copy of the input document for each element in the specified array. This flattens the nested data, allowing subsequent stages in the pipeline to access individual elements directly as if they were top-level fields of the document.

Why this answer

The $unwind stage deconstructs an array field from the input documents to output a document for each element. By using it after $lookup, it transforms an array of joined data into individual flat documents. This is a critical step when you want to group or sort by fields located inside the joined data, as those operations are much easier to perform on flat documents than on nested array structures.

Exam trap

Test-takers frequently confuse $unwind with $lookup, or forget that $unwind completely drops documents where the target array is empty unless preserveNullAndEmptyArrays is explicitly enabled.

20
MCQmedium

You are processing a large collection of user sessions. You need to calculate the average duration of sessions grouped by 'userId', but you must filter out any sessions lasting less than 5 seconds before performing the grouping. Which stage order is most efficient?

A.$group then $match
B.$project then $match then $group
C.$match then $group
D.$group then $project then $match
AnswerC

Filtering with $match at the start of the pipeline is highly efficient because it reduces the data volume before the resource-intensive $group operation. If an index exists on the duration field or the userId field, the $match stage can execute very quickly, minimizing the document count processed by subsequent pipeline stages.

Why this answer

Filtering data as early as possible in an aggregation pipeline is a fundamental optimization technique. By placing the $match stage before the $group stage, you significantly reduce the number of documents that need to be processed by the subsequent grouping operation. This reduces memory usage and improves overall pipeline performance by limiting the dataset size early on, which is critical when working with large-scale production collections.

Exam trap

Candidates often place the $group stage before the $match stage, which forces the database to process and group every document before filtering, significantly degrading performance.

21
MCQhard

An analytics collection contains order documents with an array of items. You need to output a single document for each element in the items array while preserving documents that have an empty or missing items array. Which stage configuration achieves this?

A.{ $unwind: { path: "$items", preserveNullAndEmptyArrays: true } }
B.{ $unwind: "$items" }, { $match: { items: { $exists: false } } }
C.{ $project: { items: { $ifNull: ["$items", [null]] } } }, { $unwind: "$items" }
D.{ $lookup: { from: "items", localField: "itemId", foreignField: "_id", as: "items", preserveNullAndEmptyArrays: true } }
AnswerA

Passing an options document to $unwind with path and preserveNullAndEmptyArrays set to true instructs MongoDB to emit documents even when the items array is null, missing, or empty, ensuring no parent records are silently filtered out.

Why this answer

Using the preserveNullAndEmptyArrays option within the $unwind stage ensures that parent documents lacking the target array or containing empty arrays are not dropped from the pipeline result set. This behavior is essential when performing left outer joins or maintaining complete statistical counts across documents with optional nested arrays.

Exam trap

Developers often forget that the default behavior of $unwind is to completely drop documents where the target array field is missing, null, or empty, leading to unexpected data loss.

22
MCQmedium

Which operator is most appropriate to concatenate an array of strings into a single string field?

A.$map
B.$concat
C.$reduce
D.$filter
AnswerC

$reduce allows you to iterate over an array and build a result. By using $concat as the accumulation expression, you can effectively join an array of strings into a single string. It is the most robust way to perform array-to-scalar transformations within the aggregation pipeline.

Why this answer

The $reduce operator provides a way to iterate through an array and apply an expression to each element, allowing you to build a cumulative result. When used with $concat, it is the standard and most performant way to join array elements into a single string. This is common when normalizing document data for export or display purposes in external applications.

Exam trap

Candidates often attempt to use $concat on an array directly, which causes an error because $concat expects individual string arguments, not an array of strings.

23
MCQmedium

In an aggregation pipeline, what is the purpose of the $addFields stage?

A.It removes all existing fields and replaces them with new ones.
B.It adds new fields to documents, preserving existing fields.
C.It filters the collection based on newly added fields.
D.It is a deprecated stage replaced by $project.
AnswerB

$addFields adds specified fields to the input documents, and if a field already exists, it updates its value with the result of the expression. This is the standard mechanism for augmenting documents without losing pre-existing data, providing maximum flexibility for data reporting and transformation tasks.

Why this answer

The $addFields stage allows you to append new fields to documents by defining expressions. It is highly versatile, enabling computed fields, data restructuring, or field renaming without removing existing data. This is crucial for enriching data on the fly within the database, which avoids the overhead of modifying large collections on disk or performing complex logic within the application code.

Exam trap

Candidates frequently confuse $addFields with $project, mistakenly believing that $addFields removes existing fields from the document, which is incorrect as $addFields explicitly preserves all original document fields.

24
MCQhard

You are building an aggregation pipeline that processes millions of documents. The pipeline includes a $match stage, a $group stage that calculates sums, and a $sort stage. You notice the pipeline is spilling to disk and running slowly. Which stage should you place first to optimize performance while maintaining correct results?

A.$group
B.$sort
C.$match
D.$project
AnswerC

Placing $match first reduces the number of documents flowing into subsequent stages. This leverages indexes, minimizes memory usage in $group and $sort, and prevents disk spills. It is the standard optimization for aggregation pipelines, ensuring only relevant data is processed.

Why this answer

The $match stage should be placed at the beginning of the pipeline to filter documents early, using indexes when possible. This reduces the volume of data passed to $group and $sort, lowering memory consumption and avoiding disk spills. It also maintains the correct aggregation logic by filtering before grouping and sorting.

Exam trap

The trap here is thinking that $project or $sort first might help, but they don't reduce document count or leverage indexes as effectively as $match.

25
MCQhard

What happens if a $sort stage exceeds the 100MB RAM limit without an index?

A.It automatically uses disk storage.
B.The pipeline fails with an error.
C.The sort is truncated to 100MB.
D.The query planner adds an index.
AnswerB

Without the 'allowDiskUse' option, the aggregation framework enforces a strict 100MB limit per stage to protect the system's memory. If a sort operation requires more memory, it aborts the process and throws an error, forcing the user to optimize the pipeline or enable disk usage.

Why this answer

If an aggregation stage exceeds the memory limit, the operation will fail unless 'allowDiskUse' is explicitly set to true. When enabled, MongoDB writes the data to temporary files on disk to complete the operation. While this allows large sorts to finish, it significantly degrades performance due to disk I/O, which is orders of magnitude slower than in-memory processing, highlighting the importance of proper indexing.

Exam trap

Many candidates assume MongoDB automatically writes large sort operations to disk, forgetting that the 100MB limit triggers an error unless allowDiskUse is enabled.

26
MCQeasy

An analyst needs a pipeline that reads from a `sales` collection and, for each document, emits one output document per entry in the `regions` array, copying all other fields unchanged. Which stage accomplishes this?

A.{ $unwind: "$regions" }
B.{ $group: { _id: "$regions" } }
C.{ $project: { regions: 1 } }
D.{ $replaceRoot: { newRoot: "$regions" } }
AnswerA

$unwind deconstructs an array field so that the stage outputs one document per array element, duplicating the remaining fields for each. Pointing it at the regions path produces exactly one document per region while preserving every other field, which matches the requirement without additional stages.

Why this answer

Expanding an array into individual documents is the job of $unwind, which emits one document per element and replicates the non-array fields. It requires an array path; other stages either reduce document count, replace the whole document, or merely reshape fields without altering cardinality.

Exam trap

The trap here is confusing a stage that changes document shape, such as a projection, with one that changes document count, such as $unwind.

27
MCQhard

An aggregation pipeline processes a collection of sensor readings with fields `sensorId`, `timestamp`, and `value`. The pipeline must output, for each sensor, the most recent reading's value and the average value across all readings. Which pipeline correctly achieves this?

A.$sort by timestamp descending, then $group with { _id: "$sensorId", latestValue: { $first: "$value" }, avgValue: { $avg: "$value" } }
B.$group with { _id: "$sensorId", latestValue: { $max: "$value" }, avgValue: { $avg: "$value" } }
C.$group with { _id: "$sensorId", latestValue: { $last: "$value" }, avgValue: { $avg: "$value" } }
D.$sort by timestamp ascending, then $group with { _id: "$sensorId", latestValue: { $last: "$value" }, avgValue: { $avg: "$value" } }
AnswerA

Sorting by timestamp descending ensures that within each sensor group, the first document encountered by $first is the most recent reading. The $group stage then uses $first to capture that value and $avg to compute the average. This pipeline correctly produces the desired output for each sensor.

Why this answer

To get the most recent reading's value per sensor, the pipeline must sort by timestamp descending so that the first document per group is the latest. Then $group uses $first to extract that value and $avg to compute the average. This combination ensures both requirements are met correctly.

Exam trap

The trap here is assuming $max or $last without sorting returns the latest value, when they depend on value magnitude or document order.

28
MCQmedium

What is the purpose of the $merge stage in an aggregation pipeline?

A.To filter the pipeline based on multiple conditions.
B.To replace the entire contents of a collection with new results.
C.To write results into a collection, optionally updating existing documents.
D.To sort the final output before it is returned to the client.
AnswerC

$merge allows for complex output scenarios, including updating existing documents (upserting) and inserting new documents. This makes it a powerful tool for incremental data processing and maintaining synchronized aggregate collections without needing to drop and recreate the entire target collection during every pipeline execution cycle.

Why this answer

The $merge stage is used to output the aggregation result to a collection, allowing for upserts and merging data into an existing collection. Unlike $out, which replaces the entire collection, $merge provides greater flexibility for incremental updates. This makes it ideal for maintaining materialized views or updating existing summary tables with new data without the performance cost of rewriting the entire target collection from scratch.

Exam trap

Candidates confuse $merge with $out, failing to realize that $out overwrites the entire target collection, whereas $merge allows for incremental updates and upserts into existing data.

29
MCQeasy

Which operator is used to calculate the average value of a numeric field in a $group stage?

A.$mean
B.$sum
C.$avg
D.$median
AnswerC

$avg is the standard accumulator used to calculate the arithmetic mean. It automatically handles the summation and division by the count of documents, providing a clean and efficient way to summarize numeric trends within the aggregation pipeline's grouping phase.

Why this answer

The $avg accumulator operator is used to compute the mean of all numerical values for a specific field across a group of documents. It is a fundamental tool for data analysis and reporting. Knowing how to use accumulators within the $group stage is essential for generating meaningful insights from raw data, which is a primary use case for the MongoDB Aggregation Framework.

Exam trap

Candidates sometimes confuse $avg with $sum or attempt to use it as a query operator rather than an accumulator specifically designed for use within a $group stage.

30
MCQmedium

Which operator would you use to filter for documents where an array field contains ALL of the elements in a specified array?

A.$elemMatch
B.$in
C.$all
D.$size
AnswerC

$all requires that the specified array field contains all elements specified in the query array, regardless of order. This is the correct operator for ensuring that a collection of mandatory tags or values is present, which is common in content management and inventory applications.

Why this answer

The $all operator is the standard solution for matching arrays that contain a set of required values. It is highly effective for filtering documents based on multiple criteria within a single array field, such as tags or categories. Mastery of array query operators is vital for dealing with flexible schema designs, which are a hallmark of MongoDB's document-oriented model.

Exam trap

Candidates frequently confuse $all with $in, failing to realize that $in matches documents where the field matches any value in the array, while $all requires all specified values to be present.

31
Multi-Selectmedium

A logistics company stores shipment documents in a collection named shipments. Each document has fields: `origin` (string), `destination` (string), `weight` (number), and `status` (string). The analytics team wants to find the total weight of shipments for each unique pair of origin and destination, but only for shipments with status 'delivered'. They also want to include the average weight per pair and sort the results by total weight descending. Which two stages are essential to achieve the grouping and calculation of these metrics? (Choose two.)

Select 2 answers
A.$match
B.$group
C.$limit
D.$project
E.$sort
AnswersA, B

$match is essential to filter shipments with status 'delivered' before grouping. This reduces the number of documents processed by $group and ensures that only delivered shipments are included in the totals and averages. Placing $match early in the pipeline also allows MongoDB to use an index on status if available, improving performance. It is necessary to meet the filtering condition.

Why this answer

The $match stage filters shipments to only those with status 'delivered', and the $group stage aggregates by origin and destination, computing total weight and average weight. These two stages are essential to achieve the required grouping and calculations. Other stages like $sort are used for ordering but are not essential for the core aggregation.

Exam trap

The trap here is selecting $sort as essential because the requirement mentions sorting, but the question specifically asks for stages essential to grouping and calculation.

32
Multi-Selecthard

You are reviewing an aggregation pipeline on an `orders` collection that must return only orders with a `total` greater than 100 and then compute the number of orders per `customerId`. The collection has an index on `total`. Which two stages should you use, and in what order, to maximize performance while returning the correct results? (Choose two.)

Select 2 answers
A.Place a `$group` stage with `_id: "$customerId"` and `count: { $sum: 1 }` after the filtering stage.
B.Place a `$project` stage that includes only `customerId` and `total` before the `$group` stage.
C.Place a `$match` stage with `{ total: { $gt: 100 } }` as the first stage of the pipeline.
D.Place a `$sort` stage on `total` as the first stage to speed up the filter.
E.Place a `$limit` stage of 100 before the `$group` stage to reduce the number of documents processed.
AnswersA, C

The `$group` stage with `_id: "$customerId"` and `$sum: 1` computes the number of orders per customer. Placing it after `$match` means it only processes the filtered subset, which is more efficient. This stage produces the required per-customer counts and cannot be replaced by a simple query projection because the aggregation is needed.

Why this answer

To maximize performance while returning correct counts, the pipeline should start with a `$match` on `total` so the indexed field filters documents early, then use `$group` with `$sum: 1` keyed by `customerId`. This order reduces the number of documents entering the grouping stage and ensures the counts reflect only qualifying orders. Other stages either fail to filter correctly or change the result set.

Exam trap

The trap here is thinking that any early stage reduces work, when stages like `$limit` or `$sort` can change results or add cost without using the index for filtering.

33
MCQeasy

Which stage would you use to limit the output of your aggregation to the top 10 most expensive items?

A.$limit
B.$sort followed by $limit
C.$match followed by $limit
D.$group followed by $limit
AnswerB

This sequence is correct. The $sort stage orders the documents based on the specified criteria, and the subsequent $limit stage constrains the output to the required count. This combination is the standard approach for implementing 'top N' style queries in MongoDB aggregation pipelines, ensuring accuracy in the results.

Why this answer

To achieve this, you must first sort the data by price in descending order using the $sort stage, and then use the $limit stage to capture only the first 10 documents. This is a common pattern in reporting. Getting the order of these stages correct is vital, as limiting before sorting would produce arbitrary results rather than the correct top 10 based on value.

Exam trap

Candidates frequently reverse the pipeline order, placing $limit before $sort, which results in arbitrarily truncated documents rather than the true top results.

34
Multi-Selecthard

You are troubleshooting a pipeline over a large `telemetry` collection that exceeds the aggregation memory limit while running on a single mongod instance. You must let the pipeline complete successfully without changing the server-wide configuration. Which two actions will allow the pipeline to keep running? (Choose two.)

Select 2 answers
A.Move the memory-intensive stage to the end of the pipeline.
B.Disable the memory limit by restarting mongod with a larger wiredTigerCacheSizeGB value.
C.Add a $match stage before the memory-intensive stage to reduce the document count.
D.Set the internalQueryMaxBlockingSortMemoryUsageBytes parameter to a larger value.
E.Pass the allowDiskUse option when executing the aggregation.
AnswersC, E

Filtering earlier shrinks the working set that downstream blocking stages must hold, often bringing the pipeline back under the memory threshold. Placing the $match first also lets it use indexes, so it is a legitimate way to avoid exceeding the limit without any server-wide change.

Why this answer

Two levers work within the stated constraints: enabling allowDiskUse so blocking stages can spill to disk, and shrinking the input with an early $match that can also exploit indexes. Together they address both the spill capacity and the underlying data volume, letting the pipeline complete on a single instance without any server-wide reconfiguration.

Exam trap

The trap here is reaching for server-wide memory parameters, which the scenario forbids, instead of the per-operation allowDiskUse option and early filtering.

35
MCQhard

Refer to the exhibit. What is the expected behavior if the 'price' field exists but the 'tax' field is missing in a document?

A.It treats the missing field as 0 and returns the price.
B.It returns a null value for the 'result' field.
C.The operation throws an error due to the missing field.
D.It ignores the price and returns a value of 0.
AnswerB

In MongoDB arithmetic operators, if any of the operand values are null or refer to a missing field, the entire result is null. This forces developers to be explicit about how to handle missing values by using operators like $ifNull to provide default values for calculations.

Why this answer

In MongoDB, arithmetic expressions like $add treat missing fields as null. When any input in an arithmetic expression is null, the result of the entire expression is null. Consequently, the 'result' field will be set to null for those documents.

This behavior is intentional to alert developers to incomplete data and requires explicit handling, such as using $ifNull, to ensure valid calculations are produced when expected data might be missing.

Exam trap

Test-takers frequently assume that missing fields in arithmetic expressions are treated as zero, expecting a valid calculation instead of a null result.

36
MCQmedium

Which aggregation stage allows you to replace the root of the document with a specified document or field?

A.$project
B.$replaceRoot
C.$set
D.$unwind
AnswerB

$replaceRoot effectively promotes a nested document to the top level of the aggregation stream. This is the correct choice for restructuring documents by removing the original parent structure and placing a target inner document at the top level for further processing and analysis.

Why this answer

The $replaceRoot stage is used to restructure a document by promoting a nested field to the top level. This is extremely useful for flattening complex, deeply nested JSON structures retrieved from other collections or generated through aggregations. By simplifying the document schema, it makes subsequent pipeline stages cleaner and easier to manage, which is a best practice for maintaining readable and maintainable aggregation pipelines in complex applications.

Exam trap

Candidates often mix up $replaceRoot with $project or $addFields, thinking those stages can completely swap the document root structure.

37
MCQmedium

You have a collection of sensor readings where each document includes a timestamp field `recordedAt` and a temperature field `tempC`. You need to generate a report that groups readings into 15-minute intervals and calculates the average temperature per interval. Which aggregation expression should you use to create the grouping key?

A.{ $dateTrunc: { date: "$recordedAt", unit: "minute", binSize: 15 } }
B.{ $dateFromParts: { year: { $year: "$recordedAt" }, month: { $month: "$recordedAt" }, day: { $dayOfMonth: "$recordedAt" }, hour: { $hour: "$recordedAt" }, minute: { $multiply: [ { $floor: { $divide: [ { $minute: "$recordedAt" }, 15 ] } }, 15 ] } } }
C.{ $subtract: [ "$recordedAt", { $mod: [ { $minute: "$recordedAt" }, 15 ] } ] }
D.{ $dateToString: { format: "%Y-%m-%d %H:%M", date: "$recordedAt" } }
AnswerA

The $dateTrunc operator truncates a date to the specified unit and binSize, producing a date rounded down to the nearest 15-minute boundary. This creates a consistent grouping key for intervals like 10:00, 10:15, 10:30. It is the correct tool for time-bucketing in aggregation pipelines.

Why this answer

The $dateTrunc operator is designed to truncate dates to a specified unit and bin size, making it ideal for grouping into intervals. It handles all date components automatically, ensuring correct 15-minute buckets. The other options either do not round, are overly complex, or operate on partial components, leading to incorrect grouping.

Exam trap

The trap here is assuming that $dateToString or manual arithmetic can easily round to intervals, but they require extra steps and are error-prone.

38
MCQmedium

Why does using a $limit stage immediately after a $sort stage provide a performance advantage?

A.It forces the use of a clustered index.
B.It allows the engine to keep only the top results in memory.
C.It automatically creates a temporary index.
D.It disables the need for sorting entirely.
AnswerB

By pairing these stages, MongoDB implements a top-K optimization. Instead of performing a full sort on the entire dataset, the database keeps only the top N documents in memory. This reduces memory pressure and CPU usage by avoiding unnecessary sorting work for documents that will never reach the final output.

Why this answer

When $limit follows $sort, MongoDB uses a 'top-K' algorithm to maintain only the top N results in memory throughout the sort process. This prevents the database from needing to sort the entire collection in RAM, as it can discard documents that fall outside the sorted top-K range. This optimization is crucial for large datasets where sorting the full collection would exceed memory limits or significantly impact overall system throughput.

Exam trap

Candidates often incorrectly assume that $limit after a $sort stage is just a standard filtering operation, missing the specific 'top-K' optimization that MongoDB applies to conserve memory.

39
MCQmedium

What is the primary benefit of using $unionWith to combine data from two collections in a single pipeline?

A.It automatically joins documents based on a common ID.
B.It allows querying multiple collections in one pipeline.
C.It is faster than $match because it avoids indexes.
D.It replaces the need for a $group stage entirely.
AnswerB

$unionWith provides a clean, single-pipeline approach to aggregate data from multiple collections. This reduces the latency of network round-trips from the application layer to the database, as the entire data consolidation process is executed on the server, resulting in faster and more efficient query execution.

Why this answer

$unionWith allows developers to append results from one collection to another within a single aggregation pipeline. This provides a mechanism for performing multi-collection queries without the overhead of multiple application-level calls. It simplifies data retrieval for scenarios like merging similar data from sharded collections or combining logs from multiple time-based collections, which is a powerful technique for efficient data consolidation and reporting.

Exam trap

Candidates sometimes confuse $unionWith with $lookup, believing $unionWith performs a join between collections, whereas it actually performs a set union by appending results from one collection to another.

40
MCQmedium

Refer to the exhibit. Why is this pipeline considered suboptimal?

A.The $project stage is too slow.
B.The $match stage should come before $project.
C.The _id: 0 syntax is invalid.
D.$match cannot follow $project.
AnswerB

Filtering with $match should occur early to reduce the document count. By placing $match first, you take advantage of indexing to reduce the working set. Projecting afterward ensures the transformation is only applied to the filtered subset, which is far more efficient than processing all documents first.

Why this answer

The pipeline is suboptimal because it performs projection before filtering. By projecting fields first, the database processes every document in the collection. If the $match stage were moved to the beginning, the database could potentially use an index to filter documents before performing the field projection, reducing the amount of data the projection logic needs to handle, thus improving overall performance.

Exam trap

Test-takers frequently place $project before $match, failing to realize this prevents the database from efficiently utilizing indexes.

41
MCQeasy

A developer needs to calculate the total revenue per product category from a sales collection. The collection has documents with fields `category` and `amount`. Which aggregation pipeline stage should be used to compute the sum of amounts for each category?

A.$match
B.$sort
C.$project
D.$group
AnswerD

$group is the correct stage because it groups documents by a specified key, such as category, and allows accumulators like $sum to compute aggregate values per group. This directly produces the total revenue for each category, matching the requirement precisely.

Why this answer

The $group stage is designed to group documents by a key and apply accumulators like $sum to compute aggregate values. By grouping on category and summing amount, the pipeline yields total revenue per category. Other stages like $project, $match, or $sort do not perform cross-document aggregation, so they cannot produce the required totals.

Exam trap

The trap here is assuming $project can compute aggregates across documents, when it only reshapes individual documents.

42
MCQmedium

A retail company stores orders in a collection where each document has a `customerId` field and a `total` field. The analytics team needs to see, for each customer, both the sum of all order totals and a list of the individual order totals. Which aggregation pipeline stage and accumulator combination should be used?

A.$group with { _id: "$customerId", totalSum: { $sum: "$total" }, totals: { $push: "$total" } }
B.$project with { customerId: 1, totalSum: { $sum: "$total" }, totals: { $push: "$total" } }
C.$group with { _id: "$customerId", totalSum: { $sum: "$total" }, totals: { $addToSet: "$total" } }
D.$group with { _id: "$customerId", totalSum: { $sum: "$total" }, totals: { $push: "$total" } } followed by $unwind on totals
AnswerA

This $group stage groups documents by customerId and uses $sum to compute the total of all total values, while $push collects each total value into an array. Both accumulators operate on the same field but produce different results, exactly matching the requirement for a sum and a list of individual totals.

Why this answer

The $group stage groups documents by customerId, and within it, $sum calculates the total of all total fields while $push collects each total into an array. This yields one document per customer containing both the sum and the list of individual totals, meeting the analytics requirement without extra stages that would alter the output shape.

Exam trap

The trap here is confusing $push and $addToSet in a $group stage, assuming both preserve all values when only $push includes duplicates.

43
MCQmedium

A support team stores tickets in a `tickets` collection where each document has a `status` field that may be missing on legacy records. A developer runs a pipeline whose first stage is { $match: { status: { $ne: "closed" } } } and expects every non-closed ticket to appear in the results. Which set of documents does this stage pass downstream?

A.All documents, because $match only removes documents when a field is explicitly set to null.
B.Only documents where status exists and is not equal to "closed".
C.Documents where status exists and equals "closed", because $ne inverts the comparison during aggregation.
D.Documents where status is not "closed", including documents where the field is absent.
AnswerB

The $ne operator compares values, and a missing field is treated as not matching a comparison against a string. Documents lacking status therefore fail the predicate, so only documents where status is present and differs from "closed" survive. Legacy records with no status are silently excluded from the pipeline output.

Why this answer

Comparison operators in an aggregation $match evaluate against existing field values. When a field is absent, a comparison to a string such as $ne "closed" cannot be satisfied, so those documents are filtered out. To retain them, the pipeline would need an explicit $or clause testing for field absence or null alongside the inequality.

Exam trap

The trap here is expecting $ne to behave like a logical negation that also admits documents missing the field.

44
Multi-Selectmedium

A developer is building an aggregation pipeline on a collection of blog posts. Each post document contains an array field `comments`, where each comment is an object with `author` and `text`. The developer needs to produce a separate document for each comment, preserving the post's `title` and the comment's `author` and `text`. Which TWO stages are required to achieve this? (Choose two.)

Select 2 answers
A.$sort
B.$unwind
C.$match
D.$group
E.$project
AnswersB, E

$unwind deconstructs the comments array, outputting one document per array element. This is essential to create a separate document for each comment. Without it, the array remains embedded, and you cannot easily project individual comment fields as top-level fields. Thus, $unwind is required.

Why this answer

To create a separate document for each comment, the pipeline must first use $unwind to deconstruct the comments array. Then, $project is used to shape the output documents, including the post title and the comment's author and text fields. Together, these stages achieve the required transformation.

Exam trap

The trap here is thinking $group is needed to reorganize data, but it actually aggregates, not expands, documents.

45
MCQhard

You are aggregating a `logs` collection where each document has a `timestamp` and a `level` field. You need to produce a single document that contains two arrays: one with all `timestamp` values for `level: "ERROR"` and one with all `timestamp` values for `level: "WARN"`. Which pipeline should you use?

A.[ { $group: { _id: null, errorTimes: { $push: { $cond: [ { $eq: ["$level", "ERROR"] }, "$timestamp", "$$REMOVE" ] } }, warnTimes: { $push: { $cond: [ { $eq: ["$level", "WARN"] }, "$timestamp", "$$REMOVE" ] } } } } ]
B.[ { $match: { level: { $in: ["ERROR", "WARN"] } } }, { $group: { _id: "$level", times: { $push: "$timestamp" } } } ]
C.[ { $group: { _id: null, errorTimes: { $push: { $cond: [ { $eq: ["$level", "ERROR"] }, "$timestamp", null ] } }, warnTimes: { $push: { $cond: [ { $eq: ["$level", "WARN"] }, "$timestamp", null ] } } } } ]
D.[ { $group: { _id: null, errorTimes: { $addToSet: { $cond: [ { $eq: ["$level", "ERROR"] }, "$timestamp", "$$REMOVE" ] } }, warnTimes: { $addToSet: { $cond: [ { $eq: ["$level", "WARN"] }, "$timestamp", "$$REMOVE" ] } } } } ]
AnswerA

This pipeline groups all documents into a single group using `_id: null` and uses `$push` with `$cond` to conditionally add timestamps. The `$$REMOVE` variable omits values that do not match the condition, preventing nulls from being pushed. The result is two arrays containing only the desired timestamps for ERROR and WARN levels.

Why this answer

To produce a single document with two arrays, group all documents with `_id: null` and use `$push` with a conditional expression. The `$$REMOVE` variable in the false branch of `$cond` omits non-matching values so the arrays contain only the relevant timestamps. Using `null` would introduce null entries, and grouping by `level` would produce multiple documents.

Exam trap

The trap here is using `null` instead of `$$REMOVE` in a conditional `$push`, which silently inserts nulls into the array for every non-matching document.

46
MCQmedium

A developer is building an aggregation pipeline on a collection of sensor readings. The pipeline must group documents by sensorId and compute two values: the total number of readings and the timestamp of the most recent reading. Which $group accumulator pair should be used?

A.{ $sum: 1 } and { $max: "$timestamp" }
B.{ $count: 1 } and { $last: "$timestamp" }
C.{ $sum: "$timestamp" } and { $max: "$timestamp" }
D.{ $size: 1 } and { $max: "$timestamp" }
AnswerA

Using $sum with a literal 1 counts one per document, yielding the total number of readings in each group. Applying $max to the timestamp field returns the largest (most recent) timestamp value per group, which is exactly the latest reading time. Both accumulators are valid within $group and produce the two required computed fields in a single pass.

Why this answer

Counting documents in a group requires an accumulator that increments per document, and $sum with a literal 1 achieves this. To retrieve the most recent reading, $max on the timestamp field returns the highest timestamp value in each group. These two accumulators satisfy both requirements in a single $group stage without needing an extra sort.

Exam trap

The trap here is assuming $last returns the most recent timestamp, when it actually returns the last document processed in the group unless the pipeline is explicitly sorted beforehand.

47
MCQmedium

A retail company stores order documents in a collection named orders. Each document has a field `total` (numeric) and a field `status` (string). The analytics team needs to produce a report that shows, for each `status`, the total revenue and the number of orders, but only for orders where `total` is greater than 100. The report must be sorted by total revenue in descending order. Which aggregation pipeline correctly produces this report?

A.db.orders.aggregate([ { $group: { _id: "$status", totalRevenue: { $sum: "$total" }, orderCount: { $sum: 1 } } }, { $match: { total: { $gt: 100 } } }, { $sort: { totalRevenue: -1 } } ])
B.db.orders.aggregate([ { $match: { total: { $gt: 100 } } }, { $group: { _id: "$status", totalRevenue: { $sum: "$total" }, orderCount: { $sum: 1 } } }, { $sort: { totalRevenue: -1 } } ])
C.db.orders.aggregate([ { $match: { total: { $gt: 100 } } }, { $group: { _id: "$status", totalRevenue: { $sum: "$total" }, orderCount: { $count: {} } } }, { $sort: { totalRevenue: -1 } } ])
D.db.orders.aggregate([ { $match: { total: { $gt: 100 } } }, { $group: { _id: "$status", totalRevenue: { $sum: "$total" }, orderCount: { $sum: 1 } } }, { $sort: { total: 1 } } ])
AnswerB

This pipeline first filters with $match, then groups by status and computes revenue and count, then sorts by totalRevenue descending. The $match stage uses an index if available, reducing the documents that reach $group. The $group stage correctly accumulates sum of total and count of documents, and $sort orders the results as required. This is the standard and efficient way to perform grouped aggregation with filtering and sorting.

Why this answer

The correct pipeline filters orders with total greater than 100 using $match before grouping. Then it groups by status, calculates total revenue with $sum on the total field, and counts orders with $sum: 1. Finally, it sorts the grouped results by totalRevenue in descending order.

This order of stages is efficient and produces the required report.

Exam trap

The trap here is assuming that $match can be placed after $group to filter on original fields, but after grouping the original fields are no longer present.

48
MCQmedium

You are building an aggregation pipeline on a collection of sensor readings where each document has a `timestamp` field and a `value` field. You need to compute, for each calendar day, the highest `value` and the average `value` across all readings for that day. Which `$group` stage accomplishes this?

A.{ $group: { _id: { $dateToString: { format: "%Y-%m-%d", date: "$timestamp" } }, maxValue: { $max: "$value" }, avgValue: { $sum: "$value" } } }
B.{ $group: { _id: "$timestamp", maxValue: { $max: "$value" }, avgValue: { $avg: "$value" } } }
C.{ $group: { _id: { $dateToString: { format: "%Y-%m-%d", date: "$timestamp" } }, maxValue: { $max: "$value" }, avgValue: { $avg: "$timestamp" } } }
D.{ $group: { _id: { $dateToString: { format: "%Y-%m-%d", date: "$timestamp" } }, maxValue: { $max: "$value" }, avgValue: { $avg: "$value" } } }
AnswerD

This stage groups by the formatted date string derived from `timestamp` and applies the `$max` and `$avg` accumulators to `value`. `$max` returns the highest value per group and `$avg` returns the mean, both valid accumulators. The `_id` expression using `$dateToString` correctly produces one group per calendar day, which is exactly what the scenario requires.

Why this answer

The `$group` stage must group by a day-level key while applying the correct accumulators to the `value` field. Using `$dateToString` in `_id` collapses all readings for a calendar day into one document, and `$max` plus `$avg` return the highest and mean values. Grouping by raw timestamps or using `$sum` fails to satisfy the daily maximum and average requirement.

Exam trap

The trap here is assuming that grouping by the raw `timestamp` field automatically produces daily buckets, when in fact it creates one group per distinct timestamp.

49
MCQmedium

Which stage should be placed first in an aggregation pipeline to optimize performance when filtering a large collection based on an indexed field?

A.$project
B.$match
C.$group
D.$sort
AnswerB

The $match stage serves as a query filter that can utilize existing indexes when placed at the beginning of a pipeline. By reducing the number of documents passed to downstream stages, it saves CPU and memory. This is the standard practice for performance optimization in MongoDB aggregation pipelines.

Why this answer

The $match stage filters documents before they pass to subsequent stages, significantly reducing the data volume processed by memory-intensive operations like $sort or $group. Placing it first allows MongoDB to utilize indexes, which minimizes disk I/O and RAM usage. Efficient pipeline construction is critical for scalability in production environments, as reducing the input set early directly impacts the latency and resource consumption of the entire aggregation operation.

Exam trap

Candidates often place $project or $group stages before $match, which forces MongoDB to process the entire collection, rendering the index useless and significantly increasing memory consumption and execution time.

Ready to test yourself?

Try a timed practice session using only Aggregation Framework questions.