C100DEV Data Modeling Practice Question
A social media application stores user profiles with an embedded `friends` array containing friend IDs. The array can grow to thousands of entries. Queries frequently need to find mutual friends between two users. Which schema design strategy best optimizes this workload in MongoDB?
⚠ Common exam trap
The trap here is assuming that embedding friend lists and using a multikey index is sufficient for mutual friend queries, when in fact set intersection operations on large arrays are inefficient without a dedicated relationship collection.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Model friendships as separate documents in a `friendships` collection with fields `user1` and `user2`, and create a compound index on both fields.
Modeling friendships as separate documents with a compound index on both user fields enables efficient indexed lookups for mutual friends. This design avoids the pitfalls of unbounded arrays and leverages MongoDB's indexing to handle large-scale social graphs, ensuring queries remain performant as the dataset grows.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Keep the embedded `friends` array and create a multikey index on it to support mutual friend queries.
Why it's wrong here
A multikey index on the `friends` array can support queries that match a single friend ID, but finding mutual friends requires intersecting two large arrays, which is not efficiently supported by a single multikey index. The index would be large and may not be used effectively for set intersection operations, leading to performance bottlenecks.
- ✗
Store the `friends` array as a GridFS bucket to bypass the 16MB document limit and query it with aggregation.
Why it's wrong here
GridFS is designed for storing large binary files, not for querying structured data like friend lists. Using GridFS would make it impossible to index or query individual friend IDs efficiently. This approach misuses the feature and would not support mutual friend queries, leading to severe performance and complexity issues.
- ✗
Use a `$lookup` aggregation to join the user's `friends` array with another user's `friends` array at query time.
Why it's wrong here
`$lookup` on embedded arrays is not supported directly; you would need to unwind the arrays first, which is computationally expensive and cannot use indexes efficiently. This approach would require scanning large arrays for each query, resulting in poor performance and high memory usage, especially as the number of friends grows.
- ✓
Model friendships as separate documents in a `friendships` collection with fields `user1` and `user2`, and create a compound index on both fields.
Why this is correct
Storing each friendship as a separate document with a compound index on `user1` and `user2` allows efficient queries to find all friends of a user and to compute mutual friends via set intersection using indexed lookups. This design avoids unbounded array growth and leverages MongoDB's indexing for fast, scalable queries on large datasets.
About these practice questions
Courseiva writes every C100DEV question from scratch — 259 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official MongoDB exam blueprint
This C100DEV practice question is part of Courseiva's free MongoDB certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the C100DEV exam.