Which THREE of the following are valid ways to limit the memory usage of an aggregation pipeline?
Trap 1: Increase the default memory limit to 5GB using setParameter.
The 100MB memory limit is a hard-coded constraint for individual stages in the pipeline. While allowDiskUse enables spillover, there is no configuration setting to arbitrarily increase the default RAM limit per stage, making this approach invalid for managing memory usage in MongoDB deployments.
Trap 2: Avoid using $unwind because it always duplicates data in memory.
$unwind is a necessary tool for data transformation and is not inherently bad for memory if used correctly. While it increases the document count, it is not the primary cause of memory exhaustion compared to blocking stages like $group and $sort. It should be used judiciously, not avoided.
- A
Use $match with an indexed field at the start.
Filtering with an indexed field significantly narrows the document set before it reaches memory-heavy stages like $group. This is the most effective way to reduce memory load, as the engine processes fewer documents and spends less time performing disk-based sorting or grouping.
- B
Use $project to include only the fields required for the aggregation.
Projecting only necessary fields reduces the average document size flowing through the pipeline. Smaller document sizes mean more documents can fit into the aggregation memory buffer, thereby reducing the likelihood of hitting the 100MB limit or needing to spill data to the temporary disk storage.
- C
Ensure $sort utilizes an index rather than performing an in-memory sort.
When $sort is supported by an index, the database retrieves documents in the requested order without needing to load them all into memory for sorting. This prevents blocking behavior and memory exhaustion, allowing the pipeline to handle much larger datasets than would be possible with in-memory sorting.
- D
Increase the default memory limit to 5GB using setParameter.
Why it fails: The 100MB memory limit is a hard-coded constraint for individual stages in the pipeline. While allowDiskUse enables spillover, there is no configuration setting to arbitrarily increase the default RAM limit per stage, making this approach invalid for managing memory usage in MongoDB deployments.
- E
Avoid using $unwind because it always duplicates data in memory.
Why it fails: $unwind is a necessary tool for data transformation and is not inherently bad for memory if used correctly. While it increases the document count, it is not the primary cause of memory exhaustion compared to blocking stages like $group and $sort. It should be used judiciously, not avoided.