Courseiva

CCNA Mongodb Data Modeling Questions

59 questions · Mongodb Data Modeling topic · All types, answers revealed

1
MCQmedium

You are designing a schema for a social application where each user document stores an array of their last 50 notifications. Notifications are frequently added, and the application only needs to display the most recent 50. Which schema design approach best supports this requirement while minimizing document growth?

A.Use the Bucket Pattern by grouping notifications into fixed-size buckets based on time.
B.Use the Subset Pattern by embedding only the most recent 50 notifications and storing older ones in a separate collection.
C.Use the Extended Reference Pattern by embedding notification counts and frequently accessed fields from the notification documents.
D.Use the Computed Pattern by precalculating and storing the total number of notifications per user.
AnswerB

The Subset Pattern is ideal when you frequently access a subset of related data. By embedding only the 50 most recent notifications and moving older ones to a separate collection, the user document remains small and queries for recent notifications are fast, directly meeting the requirement.

Why this answer

The Subset Pattern is the best fit because it keeps only the most frequently accessed data (the latest 50 notifications) embedded in the user document while offloading older notifications to another collection. This keeps the document small, avoids hitting the 16MB limit, and ensures fast access to recent notifications without unnecessary complexity.

Exam trap

The trap here is assuming that any pattern that reduces document size works, but the Subset Pattern specifically addresses the need to keep a capped, frequently accessed subset embedded while moving the rest elsewhere.

2
MCQeasy

An inventory system stores each product's supplier as an embedded subdocument containing the supplier name, address, and contact email. The supplier's address changes frequently, and updates now require scanning and updating thousands of product documents. Which data modeling change best resolves this maintenance problem?

A.Add an index on the embedded supplier.address field so updates to that field become faster.
B.Move the supplier subdocument into an array field so all supplier versions are retained for auditing.
C.Use the Schema Versioning Pattern by adding a version field to each product document and migrating documents lazily.
D.Convert the supplier subdocument to a reference by storing the supplier's ObjectId in the product document and keeping supplier details in a suppliers collection.
AnswerD

When an entity like a supplier is updated independently and shared by many parent documents, referencing avoids duplicated data. Storing only the supplier's ObjectId lets a single update to the suppliers collection propagate to all products, eliminating the mass-update problem while keeping document sizes small and consistent with MongoDB's recommendation to embed for locality but reference for independently mutated shared data.

Why this answer

Referencing is appropriate when a subdocument represents a shared entity that changes independently of its parent. Keeping supplier details in their own collection means one update fixes every product, removes duplicate mutable data, and keeps product documents lean. Indexing the embedded field, versioning the schema, or archiving supplier history all leave the core duplication and mass-update cost in place.

Exam trap

The trap here is reaching for an index to fix an update problem, when indexes only help reads and never remove duplicated embedded data.

3
Multi-Selecthard

A development team is designing a schema for a chat application where each conversation contains messages that grow continuously. Messages are read in chronological order, and the application must display the most recent 50 messages quickly while retaining the full history. Which TWO design decisions best support these requirements? (Choose two.)

Select 2 answers
A.Store each message as its own document in the conversations collection using a message-specific discriminator field.
B.Embed only the most recent 50 messages in the conversation document and store the full history in a separate collection.
C.Embed messages in the conversation document and rely on the 16MB limit to naturally cap growth.
D.Store messages in a separate collection with a reference to the conversation and an indexed timestamp field.
E.Store all messages in a single capped collection so older messages are automatically overwritten.
AnswersB, D

This is the Subset Pattern applied to chat: the conversation document holds a bounded, frequently accessed window of recent messages, so displaying the latest 50 requires no additional query. The separate collection retains every message for history and pagination. The bounded array keeps conversation documents small and predictable, avoiding the 16MB problem entirely.

Why this answer

The requirements split into two needs: fast access to recent messages and complete retention of history. Referencing messages in a dedicated collection with an indexed timestamp satisfies retention and efficient recent-message queries, while the Subset Pattern's bounded embedded window gives the fastest possible read for the latest 50. Together they keep conversation documents small and preserve the full message log without hitting the 16MB limit.

Exam trap

The trap here is assuming embedding must be all-or-nothing, when a bounded embedded subset combined with a full referenced collection satisfies both speed and retention.

4
MCQmedium

A financial analytics team stores daily stock snapshots as one document per symbol per day in a collection named 'prices'. Each document contains an array of 60 subdocuments, one per minute, with fields 'minute', 'open', 'close', and 'volume'. The team now needs to compute the hourly average 'close' for a single symbol over a single day. Which aggregation approach is most appropriate for this schema?

A.Run a $match for the symbol and day, then use $replaceRoot with the minute array to promote each element.
B.Use $group with $push to collect all 'close' values into an array, then apply $avg to the array.
C.Use $unwind on the minute array, then group by an expression that truncates 'minute' to the hour and average 'close'.
D.Create a separate collection with one document per minute and migrate the data, then aggregate over it.
AnswerC

Because the minute samples are embedded as an array, $unwind produces one document per minute so that a $group stage can bucket by hour and compute the average 'close'. This matches the Bucket pattern's intent of storing many measurements per document while still allowing per-measurement aggregation without reading other days or symbols.

Why this answer

The data is stored in buckets, so per-minute analysis requires expanding the embedded array before grouping. $unwind emits one document per array element, and a subsequent $group can bucket by a truncated hour value and compute the average 'close'. This preserves the storage benefits of embedding while still enabling fine-grained aggregation over the measurements.

Exam trap

The trap here is assuming that embedded arrays must be migrated to a flat collection before they can be aggregated on individually.

5
MCQmedium

What is the primary benefit of the 'Extended Reference Pattern'?

A.It completely eliminates the need for any referencing in the database.
B.It improves read performance by reducing the need for joins.
C.It allows documents to grow indefinitely without size limits.
D.It ensures that all data across all collections remains perfectly synchronized.
AnswerB

By embedding a few key fields from the referenced document, you can often satisfy the user's request without a join. This significantly reduces the overhead of constant $lookup calls, making the application faster and more responsive for the most frequent read operations while keeping the architecture flexible.

Why this answer

The Extended Reference Pattern involves copying a small amount of frequently accessed data from a referenced document into the primary document. This allows the application to retrieve the most critical information in a single read, avoiding a join for the majority of use cases. It optimizes for the common case while still allowing the full referenced document to be fetched only when truly needed.

Exam trap

Candidates often assume the pattern is for data integrity or normalization, missing that its primary purpose is performance optimization by minimizing application-level joins.

6
MCQhard

Refer to the exhibit. You are implementing the Subset Pattern for a user profile that tracks recent orders. As the number of orders per user grows indefinitely, which strategy prevents the 'unbounded array' anti-pattern?

A.Move all orders to a separate collection and reference them by ID in the user document.
B.Keep only the 10 most recent orders in the array and offload older orders to a separate collection.
C.Increase the maximum document size limit to accommodate the growing order array.
D.Use a capped collection for the orders array to automatically drop old orders.
AnswerB

This is the core implementation of the Subset Pattern. By limiting the array size to a manageable count, you guarantee the user document stays small and performant. Historical data remains accessible via a secondary query only when requested, optimizing for the common case of viewing recent activity.

Why this answer

The Subset Pattern is designed to keep the most relevant data in the primary document while moving the bulk of historical data to a separate collection. By storing only the most recent orders in the array, you prevent the document from exceeding the 16MB limit and ensure that high-frequency read operations remain performant, as smaller documents yield better cache utilization in the WiredTiger engine.

Exam trap

Candidates often assume that embedding all data is always better for performance. They fail to recognize that unbounded arrays cause document bloat, eventually hitting the 16MB limit and degrading write performance.

7
MCQmedium

A product catalog must support thousands of products whose attributes differ significantly: books have ISBN and page count, electronics have wattage and warranty, and clothing has size and material. Queries frequently filter on these type-specific attributes. Which schema design best supports this requirement in MongoDB?

A.Use the Schema Versioning Pattern with a 'schema_version' field and store attributes in version-specific subdocuments.
B.Store all attributes in a single generic 'attributes' array of key-value pairs without a category field.
C.Use the Polymorphic Pattern by storing all products in one collection with a 'category' field and category-specific attributes in a shared subdocument.
D.Create a separate collection for each product category and route queries based on category.
AnswerC

The Polymorphic Pattern keeps all related documents in a single collection even when their shapes differ, using a discriminator field such as 'category' to identify the variant. This allows a single query to filter across shared and type-specific attributes, and indexes can cover the fields each category uses. It matches MongoDB's flexible schema and avoids collection sprawl or application-side joins.

Why this answer

The Polymorphic Pattern is designed for collections whose documents share some fields but differ in others. A discriminator like 'category' lets the application and indexes target category-specific attributes while keeping all products queryable in one place. Splitting into multiple collections, flattening to generic key-value pairs, or misapplying schema versioning all sacrifice either query simplicity or schema clarity that this catalog needs.

Exam trap

The trap here is confusing schema versioning, which handles change over time, with polymorphism, which handles different shapes existing at the same time.

8
MCQeasy

A developer is designing a collection to store catalog items. Different item categories have different attributes: books have 'isbn' and 'pageCount', while electronics have 'wattage' and 'warrantyMonths'. All items share '_id', 'name', 'price', and 'category'. Which data modeling approach best fits this requirement?

A.Store all items in one collection and include only the attributes relevant to each item's category.
B.Create one collection per category, such as 'books' and 'electronics', and never query across them.
C.Normalize all attributes into a separate 'attributes' collection referenced by item documents.
D.Store all category-specific attributes in a single 'attributes' array of key-value pairs for every item.
AnswerA

The Polymorphic pattern keeps documents with similar shared fields in one collection while allowing category-specific fields to differ. Queries across all items, such as by name or price, stay simple, and category-specific queries filter on the shared 'category' field. This matches the scenario's mix of common and divergent attributes.

Why this answer

The Polymorphic pattern stores documents with a shared core plus category-specific fields in one collection. It supports queries across all items using common fields and queries within a category using the discriminator field. This is the natural fit when different document shapes share a meaningful set of fields and are queried together.

Exam trap

The trap here is assuming that differing document shapes force separate collections, when a shared field set is the signal for the Polymorphic pattern.

9
MCQhard

You have a collection of 'Users' and a collection of 'Groups'. Users can belong to many groups, and groups can have many users. Which modeling approach is most scalable for this many-to-many relationship?

A.Embed all group names in the user document.
B.Use a separate 'Membership' collection to map users to groups.
C.Store all user IDs in the group document.
D.Use the Attribute Pattern to list group IDs.
AnswerB

This approach is the most scalable. It avoids unbounded arrays in the primary documents and allows you to query memberships efficiently using indexes on either the user_id or the group_id. It is the standard way to handle many-to-many relationships without risking document size limits or data inconsistency issues.

Why this answer

In many-to-many relationships, storing arrays of references on both sides (e.g., group_ids in user, user_ids in group) can lead to unbounded arrays if either collection grows large. The most scalable approach is to use a separate 'Membership' collection that stores the mapping between User IDs and Group IDs. This prevents document bloat and allows you to index the relationship in both directions, ensuring optimal performance for all query types.

Exam trap

Candidates often attempt to use arrays on both sides of a many-to-many relationship, ignoring that both arrays will grow indefinitely as the application scales, violating the document size limit.

10
MCQhard

A financial reporting system reads account documents that each contain a nested array of the last 90 daily balance snapshots. Analysts run aggregations that only need the current balance and account type, but the full snapshot array is being loaded on every read. Which schema pattern most directly reduces the working set size for these read-heavy analytics queries?

A.Subset Pattern, keeping only the most recent snapshots in the main document and archiving older ones in a separate collection.
B.Schema Versioning Pattern, adding a version field so old documents can be migrated incrementally.
C.Computed Pattern, precomputing derived totals into new fields on the account document.
D.Bucket Pattern, grouping many accounts into a single document with shared metadata.
AnswerA

The Subset Pattern stores only the portion of data most frequently accessed in the main document while moving the rest to another collection. Since analysts only need current balance and type, trimming the 90-day array dramatically reduces document size and the working set, improving read performance without changing the analytics access path.

Why this answer

The Subset Pattern is specifically designed for documents with a large array where only a small portion is accessed regularly. By retaining the frequently used snapshots in the main document and relocating the rest, the document shrinks, the working set fits more easily in RAM, and read-heavy analytics touch far less data while still returning the fields analysts require.

Exam trap

The trap here is confusing patterns that reduce computation cost, such as Computed, with patterns that reduce document size, such as Subset.

11
MCQhard

A social platform stores user profiles in a 'users' collection and their posts in a 'posts' collection. Each post document embeds a small 'author' subdocument with 'userId', 'displayName', and 'avatarUrl'. A user changes their display name. The application must keep existing posts showing the new name. What is the most accurate statement about this Extended Reference design?

A.The design is invalid because embedding an author subdocument violates the rule that referenced data must never be duplicated.
B.The design guarantees strong consistency between the user document and all posts as long as both writes occur in a single transaction.
C.No propagation is needed because MongoDB automatically synchronizes duplicated fields across collections when the source document changes.
D.The change requires an update to the user document plus a multi-document update to every post embedding that author's duplicated fields.
AnswerD

Extended Reference duplicates a few frequently read fields, so the source of truth and the copies can diverge. When the display name changes, the application must update the user document and also propagate the new value to all post documents that embedded it. This is the known trade-off of the pattern: faster reads in exchange for write-time duplication handling.

Why this answer

Extended Reference embeds a small set of frequently accessed fields to avoid extra reads, accepting duplication as the cost. Because MongoDB does not synchronize fields across collections, any change to a duplicated field must be written to the source document and to every embedding document. The design is valid and useful, but the application owns consistency.

Exam trap

The trap here is believing that MongoDB maintains referential integrity between a source document and duplicated fields in other collections.

12
Multi-Selecthard

You have a document with an unbounded array of data. Which TWO strategies help mitigate the risk of exceeding the 16MB document size limit?

Select 2 answers
A.Implement the subset pattern to store only the most recent items in the document.
B.Increase the maximum BSON document size limit in the mongod configuration file.
C.Use the bucketing pattern to group array items into multiple documents.
D.Convert the array to a single large string field to save space.
E.Disable indexing on all array fields to prevent size expansion.
AnswersA, C

The subset pattern limits document size by storing only the most frequently accessed data inside the main document. Older or less relevant data is offloaded to a separate collection. This keeps the primary document compact, ensures faster scans, and prevents the 16MB document size limit from being reached.

Why this answer

Managing document growth is critical for long-term stability. The 'subset pattern' keeps only the most relevant items embedded, while 'bucketing' groups related items into separate documents to prevent a single document from growing indefinitely. Both techniques ensure that documents remain performant and well below the 16MB threshold, preventing application crashes when arrays grow beyond predicted sizes over time.

Exam trap

Candidates tend to choose sharding as a solution for unbounded arrays within a single document, missing that sharding operates at the collection level, not the document level.

13
Multi-Selecthard

A team is modeling a schema for an order-processing system that must support atomic updates across an order and its line items and must guarantee that reading an order never returns a partially updated set of line items. (Choose two.)

Select 2 answers
A.Store line items in a separate collection and use a multi-document transaction spanning the order and its items.
B.Enable majority read concern on the order collection, which by itself guarantees that line-item updates are applied atomically.
C.Embed line items as an array within the order document so a single-document write updates the order and its items atomically.
D.Store line items in a separate collection and rely on the order's '_id' reference to enforce atomic updates across documents.
E.Use a capped collection for line items so older writes are overwritten and consistency is maintained.
AnswersA, C

When line items must live in their own collection, a multi-document transaction provides all-or-nothing writes across the order and item documents. It also gives snapshot isolation so concurrent readers do not observe a half-applied update. This meets the atomicity and no-partial-read requirements at the cost of transaction overhead.

Why this answer

Atomic updates and no-partial-read guarantees come from either keeping related data in one document, where MongoDB's single-document atomicity applies, or wrapping separate documents in a multi-document transaction with snapshot isolation. References and read concern settings do not by themselves provide cross-document atomicity, and capped collections are unrelated to the requirement.

Exam trap

The trap here is treating a reference field or a read concern setting as if it provided atomicity across documents.

14
MCQmedium

An IoT platform ingests sensor readings every second from thousands of devices. The team wants to minimize the number of documents and index entries while still querying by device and time range. Which schema design best matches these requirements?

A.Group readings into time-bucketed documents, each holding many readings for a device over a fixed interval.
B.Store one document per reading with fields for deviceId, timestamp, and value.
C.Store each reading as a separate document and rely on a compound index on deviceId and timestamp.
D.Store one document per device containing an array of all readings since deployment.
AnswerA

Time-bucketed documents combine many readings into one document, cutting the number of documents and index entries while preserving queryability by device and time range. The bucket size is bounded, so documents stay well under 16MB. This directly addresses the stated goals of fewer documents, fewer index entries, and efficient range queries.

Why this answer

Bucketing readings into fixed-interval documents per device reduces document count and index footprint while keeping queries by device and time range efficient. Because each bucket is bounded, documents remain small and writes spread across buckets over time. The alternatives either multiply documents or create unbounded arrays that risk the size limit.

Exam trap

The trap here is treating an index as a way to reduce document count, when indexes add entries rather than consolidating data.

15
MCQmedium

Which approach is best for handling a 'One-to-Squillions' relationship in MongoDB?

A.Embed all child documents inside the parent document.
B.Store an array of child IDs in the parent document.
C.Store the parent ID in the child documents.
D.Use a separate collection for each child to ensure isolation.
AnswerC

Storing the parent ID in each child document is the standard 'referencing' approach for massive relationships. This keeps both parent and child documents small, avoids size limits, and allows you to index the parent ID on the child collection for fast, scalable retrieval of all related items.

Why this answer

For a 'One-to-Squillions' relationship, where a parent has a massive, unbounded number of children, neither embedding nor an array of references is viable. Storing the parent ID in the child documents (the 'referencing' pattern) is the only scalable solution. This allows the child documents to be queried efficiently by the parent ID without hitting document size limits or requiring massive updates to the parent document during growth.

Exam trap

Candidates often try to embed the children in the parent document, ignoring the document size limit (16MB) which makes this pattern fail for 'One-to-Squillions' relationships.

16
MCQhard

A healthcare application stores patient records. Each patient document includes a `medications` array of subdocuments, each with `name`, `dosage`, and `frequency`. The application needs to query patients who are taking a specific medication with a specific dosage. The array is expected to grow to hundreds of entries per patient. Which schema design best supports efficient querying and management of this data?

A.Embed the `medications` array as is, and create a multikey index on `medications.name` and `medications.dosage`.
B.Move medications to a separate `medications` collection with documents referencing the patient ID, and create a compound index on `{ name: 1, dosage: 1, patientId: 1 }`.
C.Use the Subset Pattern to store only the most recent 10 medications in the patient document and archive older medications in a separate collection.
D.Use the Bucket Pattern to group medications into buckets of 50 per patient document, and index the `name` and `dosage` fields.
AnswerB

Separating medications into their own collection avoids unbounded array growth in patient documents and allows for efficient querying with a compound index on `name`, `dosage`, and `patientId`. This design supports queries for specific medication and dosage across patients, and it scales well as the number of medications grows. It also simplifies updates and deletions of individual medication records.

Why this answer

Separating medications into their own collection prevents unbounded array growth in patient documents and enables efficient querying with a compound index. The index on `name`, `dosage`, and `patientId` supports queries for patients taking a specific medication with a specific dosage. This design scales well and simplifies management of individual medication records.

Exam trap

The trap here is assuming that embedding with a multikey index is always sufficient, when in fact unbounded arrays can degrade performance and complicate updates, making a separate collection more appropriate.

17
MCQeasy

A catalog stores products of many categories, each with different attributes: books have ISBN and author, while electronics have voltage and warrantyMonths. Queries always filter by category and then by category-specific attributes. Which schema approach best supports these queries?

A.Store all products in one collection with a shared set of fields and use null for attributes that do not apply.
B.Store products in one collection with a single generic attributes subdocument whose keys vary unpredictably.
C.Store each category in its own separate collection with only that category's fields.
D.Store products in one collection with a category field and category-specific fields present only where relevant, indexed appropriately.
AnswerD

A single collection with a category discriminator and only the relevant fields per product keeps documents lean and lets indexes target category-specific attributes. Queries filtering by category and then by those attributes are well supported. This matches the stated access pattern directly.

Why this answer

Using one collection with a category discriminator and only the fields relevant to each category keeps documents compact and allows targeted indexes on the attributes used in filters. Queries that filter by category and then by category-specific fields are efficient. Padding with nulls, splitting into many collections, or using unpredictable keys all undermine those queries.

Exam trap

The trap here is equating schema flexibility with unpredictable field names, when predictable per-category fields are what make indexing and filtering work.

18
MCQmedium

You are designing a schema for an e-commerce platform where products have a variable number of attributes like color, size, and material. Which modeling approach provides the best balance of flexibility and query performance for filtering products by these dynamic attributes?

A.Create a separate collection for every unique attribute type.
B.Store all attributes as top-level fields in the document.
C.Use the Attribute Pattern with an array of key-value subdocuments.
D.Serialize all product attributes into a single binary Large Object field.
AnswerC

The Attribute Pattern uses an array of subdocuments to store properties, allowing for a single multikey index. This structure effectively handles sparse data and enables efficient filtering across any attribute combination, providing high performance and query flexibility without needing to modify the collection schema when new attribute types appear.

Why this answer

Using the Attribute Pattern is the standard MongoDB practice for handling heterogeneous data fields. By restructuring attributes into an array of subdocuments containing 'k' (key) and 'v' (value) fields, you can create a single multikey index on this array. This allows the database to efficiently query across arbitrary product properties without requiring schema changes or thousands of unique collection indexes as requirements evolve over time.

Exam trap

Test-takers often create a separate index for every possible dynamic product attribute, leading to explosive index growth instead of using the Attribute Pattern.

19
MCQeasy

When modeling with MongoDB, why should you avoid 'unbounded' growth in an array field?

A.Because it makes the JSON syntax harder to read.
B.Because it can lead to hitting the 16MB document limit.
C.Because it prevents the use of secondary indexes.
D.Because MongoDB does not support arrays as a data type.
AnswerB

The 16MB limit is a hard physical constraint for every BSON document. Unbounded arrays will eventually exceed this limit, causing the database to reject further writes for those documents. This leads to critical application failures that can only be resolved by refactoring the schema to use a different pattern.

Why this answer

Unbounded arrays are a dangerous anti-pattern because every document has a hard 16MB size limit. As an array grows, the document size increases, eventually causing write operations to fail. Furthermore, updating large documents forces the database to relocate the entire object on disk, which creates significant performance overhead.

By limiting array growth, you ensure the database remains stable, performant, and within the physical constraints of the BSON format for all future operations.

Exam trap

Students often assume arrays can grow indefinitely like lists in traditional programming languages, ignoring MongoDB's rigid physical document size constraints.

20
MCQeasy

A library management system stores book information in a `books` collection. Each book document includes fields such as `title`, `author`, `ISBN`, and `genre`. The application frequently queries books by `genre` and `author`. Which index strategy is most appropriate to optimize these queries?

A.Create a compound index on `{ genre: 1, author: 1 }`.
B.Create a text index on `title`, `author`, and `genre`.
C.Create a hashed index on `genre` and a regular index on `author`.
D.Create a single-field index on `genre` and another single-field index on `author`.
AnswerA

A compound index on `genre` and `author` supports queries that filter on both fields or on `genre` alone, because the index prefix can be used. Since the application frequently queries by both `genre` and `author`, this index will efficiently serve those queries. It also supports queries that sort by `genre` and then `author`, which is common in library applications.

Why this answer

A compound index on `{ genre: 1, author: 1 }` efficiently supports queries that filter on both fields or on `genre` alone, leveraging the index prefix. It also supports sorting by `genre` and `author`, which aligns with common library queries. Single-field indexes may require index intersection, which is less efficient, while text and hashed indexes are unsuitable for exact-match queries.

Exam trap

The trap here is assuming that separate single-field indexes are equivalent to a compound index, or that a text index can handle exact-match queries, when in fact a compound index is specifically optimized for the query pattern.

21
MCQmedium

When designing a schema for a time-series dataset, why is it recommended to use the 'Bucket' pattern?

A.It allows for unlimited document growth without checking the 16MB limit.
B.It reduces index size and improves read performance by grouping data.
C.It automatically converts all data to a binary format for faster storage.
D.It enables ACID transactions for non-related collections automatically.
AnswerB

The Bucket pattern aggregates multiple individual readings into a single document. This creates a much smaller index, as there is one entry per bucket rather than one entry per individual reading. This allows the working set to fit better in memory, dramatically increasing the speed of time-based query operations.

Why this answer

The Bucket pattern groups multiple related documents (like sensor readings) into a single document representing a time window (e.g., one hour). This significantly reduces the total number of documents in the collection, which improves index efficiency and reduces the memory footprint. By minimizing the index size, the database can keep more of the index in RAM, leading to faster query performance for analytical or time-based data retrieval.

Exam trap

Candidates often confuse the Bucket pattern with simple embedding. They assume it's just about nesting data, missing the critical aspect of grouping readings by time windows to optimize index size.

22
MCQhard

Refer to the exhibit. Which design pattern is used to store 'total_price' in the document?

A.The Subset pattern.
B.The Computed pattern.
C.The Polymorphic pattern.
D.The Schema Versioning pattern.
AnswerB

The Computed pattern stores calculated results directly in the document, which speeds up read operations. By avoiding complex aggregations during queries, the application serves data much faster. This pattern is commonly used for values like totals, averages, or counts that are derived from other fields in the document.

Why this answer

This is the Computed pattern. By storing the final total in the document, the application avoids recalculating it during every read request. This is particularly valuable in heavy read-heavy applications where the cost of performing arithmetic aggregations across multiple documents is significantly higher than the cost of updating a pre-computed value during the initial write operation.

It optimizes read performance at the expense of write complexity.

Exam trap

Candidates often assume totals or aggregations should always be calculated dynamically on read, forgetting the heavy performance toll in read-heavy systems.

23
MCQmedium

Which schema pattern is demonstrated in the exhibit provided?

A.The Linking Pattern.
B.The Embedding Pattern.
C.The Polymorphic Pattern.
D.The Computed Pattern.
AnswerB

The exhibit displays a classic embedding pattern where order objects are stored inside the user's 'orders' array. This structure ensures that when the user document is retrieved, all relevant order information is fetched in the same query, reducing the need for additional lookups or complex joins in code.

Why this answer

The exhibit shows an array of sub-documents nested directly within a parent document, which is the definition of the embedding pattern. This pattern is used to keep related data together in one physical storage unit. It is highly efficient for read-heavy applications where the parent user record and their order history are almost always required at the same time by the application layer.

Exam trap

Candidates often confuse the embedding pattern with referencing because they see sub-documents, mistakenly thinking that any array of ObjectIDs represents a relationship instead of actual inline data.

24
MCQmedium

A social media platform stores user profiles in a `users` collection. Each user document includes a `followers` array of user IDs. For highly popular accounts, this array has grown to over 2 million entries, causing documents to approach the 16 MB BSON limit and slowing read operations. The application frequently displays a user's follower count and the first 20 followers. Which schema design change best addresses this issue?

A.Convert the `followers` array to a GridFS bucket to store the list of follower IDs as a binary blob.
B.Normalize the schema by moving the entire `followers` array into a separate `followers` collection with one document per follower relationship.
C.Use the Outlier Pattern: store the first 1000 followers in the user document and flag the document as an outlier; store remaining followers in a separate `followers` collection keyed by user ID.
D.Apply the Attribute Pattern by converting the `followers` array into a set of key-value pairs where each follower ID is a field name.
AnswerC

The Outlier Pattern is designed for documents that exceed typical size limits due to a small number of outliers. By embedding only a subset of followers and flagging the document, the common case remains fast while the full list is accessible via a separate collection. This directly resolves the 16 MB limit and performance degradation while preserving the ability to show the first 20 followers and count.

Why this answer

The Outlier Pattern is specifically designed to handle documents that grow beyond typical limits due to a few outliers. By embedding only a subset of followers and storing the rest separately, the design keeps the common case fast while accommodating the extreme case. This maintains the ability to quickly show the follower count and first 20 followers without hitting the 16 MB limit.

Exam trap

The trap here is assuming that any large array must be fully normalized or moved to GridFS, rather than recognizing that the Outlier Pattern preserves fast access to the common subset while isolating the overflow.

25
MCQeasy

A retail catalog team stores each product as a document. Products share core fields such as sku, name, and price, but electronics carry warrantyMonths while apparel carries sizeChart. The application queries all products uniformly by category and price. Which modeling approach best fits this requirement?

A.Store category-specific attributes in a separate attributes collection and $lookup them when rendering each product.
B.Normalize the schema by storing all category-specific fields in a shared subdocument with a fixed set of keys for every product.
C.Store all products in one collection and allow documents to carry category-specific fields alongside the shared fields.
D.Create a separate collection per product category so each collection has a uniform schema.
AnswerC

Keeping one collection with shared core fields plus optional category-specific fields lets a single indexed query on category and price serve every product. MongoDB does not require uniform document shape, so the polymorphic documents coexist cleanly, and the application handles optional fields only where relevant. This matches the uniform query requirement with the least complexity.

Why this answer

MongoDB schemas are flexible by design, so a single products collection can hold documents that share common fields and add category-specific ones as needed. This supports one uniform query path indexed on category and price, avoids joins, and lets new categories be introduced without schema migrations. Splitting collections or normalizing into fixed keys adds routing and storage costs that the scenario does not justify.

Exam trap

The trap here is treating MongoDB like a relational database and assuming each variant shape needs its own table or a fixed set of columns, when a single collection with optional fields is the intended design.

26
Multi-Selectmedium

Which TWO of the following are benefits of using the Polymorphic Pattern in MongoDB?

Select 2 answers
A.It allows a single query to retrieve all related items across different types.
B.It enforces strict schema validation for all documents within the collection.
C.It simplifies management by grouping related entities into one collection.
D.It automatically increases the write throughput of the database cluster.
E.It allows for unlimited document size by splitting fields across collections.
AnswersA, C

Because all polymorphic entities reside in the same collection, a single find operation can return all relevant items, regardless of their specific type. This avoids running multiple queries against different collections and merging results in the application, which simplifies code and reduces database round-trip times for complex operations.

Why this answer

The polymorphic pattern allows you to store different document structures in the same collection, which is useful when entities share common fields but have unique characteristics. By keeping these in one collection, you simplify queries that need to target the entire entity set and reduce the need for complex application-level logic to manage multiple collections for similar object types, improving developer productivity and consistency.

Exam trap

Candidates often assume the polymorphic pattern completely eliminates the need for sparse or partial indexes, forgetting that different document structures may require targeted indexing strategies.

27
MCQhard

Refer to the exhibit. Which MongoDB design pattern is represented here?

A.The Bucket pattern.
B.The Attribute pattern.
C.The Extended Reference pattern.
D.The Computed pattern.
AnswerB

The Attribute pattern allows for efficient indexing of various attributes within a single collection. By storing attributes in an array of key-value pairs, you can create a single multikey index that covers all possible attributes. This is vital when the number of potential attributes is large or dynamic.

Why this answer

This is the Attribute pattern. It is used when you need to store documents where fields vary significantly or when you need to index multiple fields that share a common structure. By transforming key-value pairs into an array of sub-documents, you can create a single multikey index on the 'specs.k' and 'specs.v' fields, allowing for efficient queries across diverse attribute sets that would otherwise require too many individual indexes.

Exam trap

Students frequently mistake the Attribute pattern for a standard dynamic schema with random root fields, failing to recognize its specific use of key-value subdocument arrays.

28
MCQmedium

You are designing a schema for a social media platform. A user has a 'profile' document, and they can have thousands of 'followers'. How should you model the follower relationship?

A.Embed all follower IDs in an array within the user profile document.
B.Create a separate collection for followers and store references.
C.Use the Attribute Pattern to store each follower as a key-value pair.
D.Store the followers in a gridFS file to handle the large size.
AnswerB

Storing followers in their own collection allows the system to scale to millions of followers per user. You can index the user ID field in this collection to perform fast lookups, ensuring that the user document remains small and that follow counts can be efficiently managed via aggregation queries.

Why this answer

Storing thousands of followers in an array within the user profile document is a violation of the unbounded array anti-pattern, as it would cause the document to grow beyond the 16MB limit and result in poor write performance. Referencing the followers in a separate collection, where each follower is a separate document, provides a scalable solution that supports high growth and efficient querying without performance degradation.

Exam trap

Candidates often assume that because MongoDB is document-oriented, they should embed everything. They fail to account for the 'unbounded' nature of social media followers that will eventually break the document.

29
MCQeasy

A team is designing a schema for a product catalog where different product categories have entirely different attributes: books have ISBN and page count, electronics have wattage and warranty period, and clothing has size and material. All products must be searchable in a single query by name. Which data modeling approach best fits this requirement?

A.Create a separate collection per product category and query them with $unionWith for every search.
B.Normalize all attributes into a generic key-value array so every product has identical field names.
C.Store all products in one collection, allowing each document to have its own category-specific fields alongside common fields.
D.Embed every possible attribute for all categories in each document, leaving unused fields null.
AnswerC

The Polymorphic Pattern stores documents of different shapes in one collection, sharing common fields like name and price while allowing category-specific fields. This lets a single indexed query on name retrieve all product types and is the idiomatic MongoDB approach for heterogeneous entities that share a common access path.

Why this answer

The Polymorphic Pattern is the standard MongoDB approach when multiple entity types share some fields and a common query path but differ in their specific attributes. Keeping all products in one collection with shared fields such as name and category-specific fields lets a single indexed query serve the catalog search while preserving each category's natural structure.

Exam trap

The trap here is assuming that heterogeneous documents must be split into separate collections, when MongoDB is designed to handle varied shapes within one collection.

30
MCQmedium

When designing a schema for a blog application, you need to store comments for posts. The comments grow indefinitely. Which approach is most effective for long-term scalability?

A.Embed all comments in the post document.
B.Reference all comments in a separate collection.
C.Use the Hybrid pattern for embedded and referenced comments.
D.Use the Attribute pattern for every comment.
AnswerC

The Hybrid pattern balances the speed of embedding with the scalability of referencing. By keeping the latest comments embedded, you ensure fast reads for the most common use cases, while offloading historical data to a separate collection, preventing document bloat and ensuring the database scales as expected.

Why this answer

The Hybrid pattern is best here. You embed the most recent comments (e.g., the last 10) directly within the post document to allow for fast initial page loads. Simultaneously, you store older comments in a separate collection with a reference to the post ID.

This provides a balance between high-speed access for recent content and the ability to scale indefinitely without ever hitting the BSON size limit.

Exam trap

Candidates often choose only one extreme: either embedding all comments or referencing all of them. They fail to realize that a hybrid approach is necessary for balancing performance and scalability.

31
MCQhard

An IoT platform ingests temperature readings from thousands of sensors. Each reading is currently stored as its own document with sensorId, timestamp, and value, producing millions of tiny documents per day. Queries typically retrieve readings for a sensor over a specific hour. Which schema design most improves storage efficiency and query performance for this access pattern?

A.Enable sharding on the readings collection with sensorId as the shard key and store one reading per document.
B.Add a compound index on sensorId and timestamp and continue storing one reading per document.
C.Use the Bucket Pattern, grouping readings from a sensor for a time window into a single document with an array of measurements.
D.Apply the Attribute Pattern by moving each reading's value into a key-value subdocument.
AnswerC

The Bucket Pattern consolidates many readings into one document per sensor per time window, dramatically cutting document count and per-document overhead. Since queries target a sensor over an hour, a bucket keyed by sensorId and hour aligns perfectly, letting a single document read satisfy the query and reducing index size.

Why this answer

The Bucket Pattern is purpose-built for time-series data, grouping measurements into documents keyed by a time window. Because queries retrieve a sensor's readings for a specific hour, a bucket per sensor per hour lets one document serve the query, slashes document count and overhead, and shrinks the index footprint, improving both storage efficiency and read performance.

Exam trap

The trap here is reaching for indexing or sharding to fix a problem that is fundamentally about document granularity and per-document overhead.

32
MCQmedium

A social media application stores user posts. Each post document currently embeds a comments array. A single popular post can accumulate over 50,000 comments, causing the document to grow beyond 10 MB. The development team wants to avoid hitting the 16 MB document size limit while still being able to retrieve a post with its most recent 10 comments in a single query. Which schema design should they implement?

A.Keep the most recent 10 comments embedded in the post document and store older comments in a separate comments collection.
B.Move all comments into a separate comments collection and store an array of comment _ids in the post document.
C.Use a capped collection for the posts collection to automatically limit document size.
D.Convert the comments array into a GridFS bucket and store the file reference in the post document.
AnswerA

This approach implements the Subset Pattern: by keeping only the most recent 10 comments embedded, the post document size remains bounded and well below 16 MB. The application can retrieve the post with its recent comments in a single query, while older comments are stored separately and can be fetched on demand, balancing performance and document size.

Why this answer

The Subset Pattern is ideal when a document contains a large array that grows indefinitely. By embedding only the most recent 10 comments, the post document remains small and the application can still fetch the post with its recent comments in one read. Older comments are stored separately, preserving access to historical data without bloating the main document.

Exam trap

The trap here is assuming that moving all comments to a separate collection and storing an array of _ids solves the size problem, but the array of _ids itself can still grow beyond the 16 MB limit.

33
MCQeasy

When designing a schema for a one-to-many relationship, what is the primary factor in deciding between embedding and referencing?

A.The total number of documents in the collection.
B.The access pattern and the size of the related data.
C.The specific version of MongoDB being used by the server.
D.The hardware specifications of the database server.
AnswerB

Access patterns define whether data is retrieved together, while size constraints prevent hitting the 16MB limit. Balancing these requirements determines if the data should be embedded for performance or referenced for scalability. This is the fundamental trade-off in MongoDB schema design that dictates document structure and query efficiency.

Why this answer

The primary factor is the access pattern: how often the data is read and whether the related data is needed together. Embedding is ideal when the child data is frequently accessed with the parent. Referencing is necessary when the child data is large, grows unbounded, or is accessed independently.

Choosing the right approach ensures that the application maintains optimal performance by minimizing read latency and avoiding document size limits.

Exam trap

Candidates often choose embedding strictly based on relational foreign key habits or database size limits, ignoring the application's actual read and write access frequency patterns.

34
Multi-Selectmedium

A team is deciding how to model the relationship between orders and the products each order contains. They want to optimize for the common query that retrieves an order along with the name and current price of every product in it. Which two modeling considerations best support this requirement? (Choose two.)

Select 2 answers
A.Embedding a copy of the product name and price inside each order line avoids a second lookup when rendering the order.
B.Modeling order lines as a separate collection joined by orderId avoids any duplication and is always the best choice for order reads.
C.Accepting some duplication of product fields is reasonable when those fields change infrequently and the read savings are significant.
D.Storing only the product ObjectId in each order line and resolving names and prices at read time with $lookup keeps data fully normalized.
E.Duplicating the entire product document into each order line ensures the order always reflects product details.
AnswersA, C

Embedding the small set of product fields the order view needs means a single document read returns everything required to display the order. This is the Extended Reference idea: store only the fields the order actually uses, keeping documents small while eliminating the join for the common read path.

Why this answer

The requirement is a frequent read that needs an order plus a few product fields. Embedding just those fields as an extended reference satisfies the query with one document read, and accepting limited duplication is justified because the copied fields change rarely. Together these considerations optimize the common read path without inflating documents with unnecessary product data.

Exam trap

The trap here is treating duplication as inherently wrong, when in MongoDB a small, rarely changing copy of fields can be the correct optimization for a hot read path.

35
MCQhard

When should you use the 'Computed Pattern' in MongoDB?

A.When you have a high write-to-read ratio for the specific data field.
B.When the data is accessed frequently and requires an expensive calculation.
C.When you need to store binary large objects (BLOBs) in the database.
D.When you want to prevent all indexing on a specific collection.
AnswerB

Pre-computing expensive values and storing them in the document allows for lightning-fast reads. The application avoids the performance penalty of re-calculating results (like sums or averages) during every query. This is a classic trade-off where storage space is exchanged for significant gains in read performance and application responsiveness.

Why this answer

The Computed Pattern is used to store calculated values (like totals, counts, or averages) directly within a document to minimize expensive real-time calculations. This is vital when the cost of computing a result exceeds the cost of updating a stored value. It is particularly effective for read-intensive workloads where the computed result is requested frequently, allowing the application to simply read the pre-calculated value from the document.

Exam trap

Candidates frequently suggest using the Computed Pattern for all calculations, failing to realize it is only beneficial when the cost of frequent reads outweighs the cost of occasional updates.

36
Multi-Selecthard

A social application stores posts with an embedded array of likes containing userId values. Some posts have millions of likes. The team wants to keep post documents small and still answer questions such as whether a specific user liked a post and how many likes a post has. Which two design changes best meet these goals? (Choose two.)

Select 2 answers
A.Store likes in a capped collection and read the count with countDocuments.
B.Move likes into a separate collection where each document records a postId and userId pair.
C.Keep the full likes array on the post and add an index on the array field.
D.Embed only the first 100 likes in the post and discard the rest.
E.Maintain a likeCount field on the post document that is incremented and decremented as likes change.
AnswersB, E

Storing each like as its own document removes the unbounded array from the post, keeping post documents small and well under the 16MB limit. It also allows an index on postId and userId to answer whether a given user liked a post efficiently. This directly addresses the size and lookup goals.

Why this answer

Moving likes to their own collection removes the unbounded array from posts and enables indexed per-user lookups, while a denormalized likeCount on the post answers the count question without embedding all likes. Together these keep post documents small and support both required questions. Indexing the array, capped collections, or truncation do not satisfy both goals.

Exam trap

The trap here is believing that indexing an array solves document growth, when indexes improve query speed but leave the oversized document in place.

37
Multi-Selectmedium

Which TWO of the following scenarios are best suited for using the 'Subset Pattern' in MongoDB?

Select 2 answers
A.A blog post collection where you display the last 5 comments on the main view.
B.A shopping cart collection where items are added and removed constantly.
C.A product catalog where you need to filter by over 100 different categories.
D.A user profile collection where you store their 'top 3' favorite movies.
E.An audit log collection where every action must be stored for compliance.
AnswersA, D

This is the classic use case for the Subset Pattern. By embedding only the most recent five comments in the post document, you ensure fast page loads. The remaining comments are stored in a separate collection and retrieved only when the user explicitly requests to see the full comment history.

Why this answer

The Subset Pattern is used to prevent documents from growing too large while keeping frequently accessed data local. By embedding only the most important subset of child documents (e.g., the five most recent comments) and referencing the rest, you balance read performance with document size management. This is essential for maintaining consistent performance levels as datasets scale, ensuring that the working set remains efficient for typical user queries.

Exam trap

Test-takers often apply the Subset Pattern to unbounded arrays that grow continuously without restriction, defeating the purpose of limiting document size.

38
MCQmedium

An e-commerce application models product reviews inside an array on the product document. As popular products accumulate hundreds of thousands of reviews, write operations and document fetch operations experience performance degradation. Which data modeling pattern best resolves this issue?

A.Keep the array embedded but add a partial index to speed up searches on the most recent reviews inside the array.
B.Move every review into its own separate collection and establish a manual join reference back to the parent product ID.
C.Implement the Bucket Pattern by storing a limited number of reviews per document alongside grouping metadata, creating new bucket documents as needed.
D.Convert the reviews array into a capped collection and use tailing cursors to push updates directly to the product document.
AnswerC

The Bucket Pattern caps reviews per document and spreads them across multiple bucket documents with grouping metadata. This bounds individual document size, so reads and writes no longer degrade as reviews accumulate, resolving the unbounded array growth.

Why this answer

Migrating from an unbounded embedded array to the Bucket Pattern divides reviews into manageable document chunks based on a count or time threshold. This prevents documents from exceeding the 16MB limit and reduces RAM pressure. Managing growth ensures predictable update performance and maintains indexing efficiency across high-throughput collections in production workloads.

Exam trap

Candidates often assume embedding is always superior in MongoDB, forgetting that arrays growing without a predictable upper bound eventually violate the 16MB document size limit and degrade write performance.

39
MCQmedium

A social media platform stores each user's 'followers' as an array of ObjectIds inside the user document. Analytics shows that a small number of celebrity accounts have millions of followers, and these documents are approaching the 16MB BSON limit. Which schema pattern should a MongoDB developer apply to resolve this issue?

A.Apply the Bucket Pattern by grouping follower ObjectIds into fixed-size buckets of 1000 per document in a separate collection.
B.Apply the Outlier Pattern by adding an overflow flag and moving the followers array into a separate collection only for users that exceed a threshold.
C.Apply the Extended Reference Pattern by duplicating the ten most important follower documents into the user document.
D.Apply the Subset Pattern by storing only the 100 most recent followers in the user document and the full list in a separate collection.
AnswerB

The Outlier Pattern handles exactly this case: most documents fit comfortably in the common shape, but a few outlier documents would exceed limits. By flagging outliers (for example with a 'has_extra' boolean) and storing the overflow in a companion collection, the schema stays efficient for the majority while celebrities are handled correctly without redesigning the whole model.

Why this answer

The Outlier Pattern is the correct design when a minority of documents drastically exceed the typical size of the rest. By marking outlier documents and relocating their oversized arrays to a separate collection, developers preserve the simple embedded model for the common case while still supporting celebrity-scale data. Neither subsetting recent entries nor time-bucketing events addresses the real requirement of storing a complete, queryable follower list for every user.

Exam trap

The trap here is assuming any large array must be subsetted, when the schema only breaks for a small number of outlier documents.

40
MCQmedium

A logistics company uses MongoDB to store shipment data. Each shipment document includes a status field that can be one of several values: 'pending', 'in_transit', 'delivered', 'cancelled'. The application frequently queries shipments by status and also needs to generate reports that count shipments per status. The team wants to ensure efficient queries and minimal index overhead. Which schema design consideration is most important?

A.Use a compound index on status and shipment_date to support both filtering and sorting.
B.Store the status as an integer code and create a single-field index on it.
C.Use the Attribute Pattern to store status as a key-value pair and index the key and value fields.
D.Store the status as a string and create a single-field index on it.
AnswerD

Storing the status as a string and indexing it is the most straightforward and efficient approach for this scenario. A single-field index on the status field supports equality queries and can be used for counting via aggregation. It has minimal overhead and is well-suited for low-cardinality fields, though in this case the cardinality is moderate. This design is simple and aligns with MongoDB best practices.

Why this answer

For a field with a limited set of values that is frequently queried, a simple single-field index on that field is the most efficient and straightforward solution. It supports equality matches and aggregation for counting, and it has minimal impact on write performance. More complex patterns or compound indexes are not warranted based on the described access patterns.

Exam trap

The trap here is over-engineering the schema by using patterns designed for variable fields or adding unnecessary compound indexes when a simple index suffices.

41
MCQhard

A financial application stores account documents with an embedded array of transactions. The array is unbounded and documents have reached 8MB. Auditors require a complete, ordered history of every transaction. Which approach best preserves full history while keeping account documents within size limits?

A.Move transactions to a separate collection, each transaction referencing its account, and use the Bucket Pattern to group transactions by month for efficient range queries.
B.Embed transactions but compress them using BSON binary subtype 0 so they consume less space.
C.Keep transactions embedded but cap the array at 1000 entries, discarding the oldest when new ones arrive.
D.Store transactions in a GridFS bucket because GridFS is designed for data that exceeds 16MB.
AnswerA

Separating transactions into their own collection removes the unbounded growth from the account document, and bucketing them by month keeps related transactions physically close and index-friendly for the date-range queries auditors run. Each transaction still references its account, so the complete ordered history is preserved, and no single document approaches the 16MB limit even for very active accounts.

Why this answer

The correct design separates the unbounded child data into its own collection and uses bucketing to keep related transactions efficient to query. Referencing the account preserves relational integrity, while monthly buckets cut index size and improve range-scan performance. Capping, compressing, or pushing structured data into GridFS all fail either the completeness requirement or the queryability requirement that auditors demand.

Exam trap

The trap here is treating the 16MB limit as the only constraint and ignoring the audit requirement that every transaction must remain queryable and ordered.

42
MCQmedium

An e-commerce application stores orders in an `orders` collection. Each order document contains an array of `items`, with each item having `productId`, `quantity`, and `price`. The application needs to frequently generate reports that sum the total revenue per product across all orders. Which schema design or feature best supports this requirement efficiently?

A.Use the Bucket Pattern to group orders by product ID into buckets, then sum the revenue within each bucket.
B.Use the Extended Reference Pattern to embed product details into each order item, then use aggregation to sum revenue.
C.Use the Computed Pattern to store a `totalRevenue` field in each product document, updating it whenever an order is placed.
D.Use the Outlier Pattern to store high-volume product orders separately, then aggregate across both collections.
AnswerC

The Computed Pattern precomputes and stores aggregated values, such as total revenue per product, directly in the product document. This eliminates the need to scan all orders and aggregate at query time, making report generation fast and efficient. Updates occur on order placement, which balances write overhead with read performance. This is ideal for frequently accessed aggregates.

Why this answer

The Computed Pattern is designed to precompute and store aggregated values like total revenue per product. By updating the product document whenever an order is placed, the application avoids expensive aggregation scans during report generation. This pattern is ideal for frequently accessed aggregates and provides fast read performance at the cost of additional write operations.

Exam trap

The trap here is assuming that aggregation pipelines alone are sufficient for frequent reporting, when in fact precomputing aggregates with the Computed Pattern is more efficient for read-heavy workloads.

43
MCQeasy

A mobile banking app stores each customer's account document with an embedded array of the last 20 transactions for quick display, while the full transaction history lives in a separate transactions collection. Which data modeling consideration most directly justifies this split?

A.The separate collection enforces referential integrity on the transaction documents.
B.Storing history separately allows the account document to exceed 16MB without error.
C.The embedded array keeps the frequently accessed working set small while the unbounded history is stored separately.
D.Embedding transactions guarantees ACID transactions across the account and history collections.
AnswerC

Embedding only the most-recently-used transactions keeps the account document compact, so the working set fits in RAM and the 16MB document limit is not threatened by unbounded growth. The complete history is preserved in a dedicated collection that can be queried when needed. This is the classic rationale for splitting high-frequency small data from large, growing data.

Why this answer

Keeping a small, bounded embedded array of recent items alongside an unbounded separate collection is a common way to keep the hot working set small and avoid unbounded document growth. The account document stays compact and fast to read, while the full history remains queryable. The other choices describe features MongoDB does not provide through embedding.

Exam trap

The trap here is assuming that embedding data provides cross-collection ACID guarantees or referential integrity, when MongoDB provides neither automatically.

44
MCQmedium

You are modeling an IoT telemetry system in MongoDB. Each sensor device emits a reading every 5 seconds, and the application most often queries the last 24 hours of readings for a single device. You want to minimize the number of index entries and document reads per query. Which schema design should you choose?

A.Store readings in a separate collection and use $lookup to join them to a device document at query time.
B.Store readings in per-device, per-hour bucket documents that hold a readings array, with a compound index on { deviceId: 1, startTime: -1 }.
C.Store each reading as its own document with a compound index on { deviceId: 1, timestamp: -1 }.
D.Store all readings for a device in a single document containing a readings array that grows indefinitely.
AnswerB

Bucketing collapses 720 readings per hour into one document, so a 24-hour query touches roughly 24 documents and 24 index keys instead of tens of thousands. This directly reduces both index entries and document reads, and the { deviceId: 1, startTime: -1 } index efficiently bounds the time range for a single device.

Why this answer

Bucketing groups many measurements into a single document, so a fixed time-window query for one device scans a small, predictable number of documents and index keys. Grouping by device and hour keeps buckets bounded in size, avoiding the 16MB limit, while the index on deviceId and startTime supports fast range filtering. The other designs either explode the number of documents or introduce unbounded growth and joins.

Exam trap

The trap here is assuming that one document per reading is the most granular and therefore most flexible design, when in fact it multiplies index entries and read operations for the fixed time-window queries this workload actually performs.

45
Multi-Selecthard

A development team is designing a schema for an e-commerce application that stores product data. Products have a set of core fields (name, price, description) that are common to all products, but different product categories have unique attributes (e.g., 'screen_size' for electronics, 'fabric' for clothing). Queries often filter on these category-specific attributes. The team wants a flexible schema that supports efficient queries on any attribute. Which TWO of the following design approaches are most appropriate? (Choose two.)

Select 2 answers
A.Use the Bucket Pattern to group products by category into buckets of 100 documents.
B.Embed all possible category-specific attributes in every product document, using null for missing attributes.
C.Use the Polymorphic Pattern to store all products in one collection with a 'category' field and include category-specific fields at the top level.
D.Use the Attribute Pattern to store category-specific attributes as an array of key-value pairs, and create a compound index on the keys and values.
E.Create a separate collection for each product category with its own schema and indexes.
AnswersC, D

The Polymorphic Pattern allows documents in the same collection to have different shapes, which is perfect for products with varying attributes. By including a 'category' field, the application can easily filter by category and access the specific attributes. This approach provides flexibility and is a common MongoDB design pattern for heterogeneous data, enabling efficient queries on category-specific fields when indexed appropriately.

Why this answer

The Polymorphic Pattern allows a single collection to store documents with different structures, while the Attribute Pattern converts variable attributes into an indexable array of key-value pairs. Together, they provide a flexible schema that supports efficient queries on any category-specific attribute. Both patterns are widely used in MongoDB for handling heterogeneous data and are appropriate here.

Exam trap

The trap here is assuming that separate collections per category or embedding all possible fields are good solutions, but they lead to complexity or inefficiency; the Polymorphic and Attribute patterns are the correct choices for flexible, queryable schemas.

46
MCQmedium

What happens if a document grows beyond the 16MB limit due to an 'Embedding' design choice?

A.MongoDB automatically moves the overflow data to a new document.
B.The write operation will fail.
C.The engine automatically compresses the document to fit.
D.The document is split into two, and the original ID is shared.
AnswerB

The 16MB limit is a hard architectural constraint. If an update causes a document to exceed this size, the database rejects the change and returns an error. This is why careful schema modeling, such as using the Subset pattern or Referencing, is essential for datasets that grow over time.

Why this answer

When a document exceeds 16MB, the MongoDB write operation fails, and the database returns an error. The BSON document size limit is a hard constraint in MongoDB. This scenario highlights why 'unbounded array' design patterns must be avoided.

When designing schemas, developers must ensure that embedded data sets have a predictable upper bound, or they must switch to a referencing strategy to maintain the integrity of their data storage.

Exam trap

Many students think MongoDB automatically truncates oversized data or splits documents across pages, forgetting that the 16MB limit is a strict rule causing total write failure.

47
MCQeasy

A team is designing a schema for a MongoDB application that stores user profiles. Each profile includes a list of the user's favorite movies, which is typically fewer than 20 items and updated infrequently. Queries often retrieve the entire profile along with the favorite movies. Which data modeling approach is most appropriate?

A.Normalize the data by creating a separate collection for movies and a junction collection linking users and movies.
B.Use the Bucket Pattern to group favorite movies into buckets of 10.
C.Store favorite movies in a separate collection and reference them by _id in the user profile.
D.Embed the favorite movies as an array of subdocuments within the user profile document.
AnswerD

Embedding the favorite movies array within the user profile document is ideal here because the list is small, updated infrequently, and always retrieved together with the profile. This approach provides atomic updates and fast reads with a single query, avoiding the need for joins. It aligns with MongoDB's recommendation to embed when there is a one-to-few relationship and data is accessed together.

Why this answer

For a one-to-few relationship where the embedded data is small, updated infrequently, and always accessed with the parent document, embedding is the recommended approach. It simplifies queries, ensures atomicity, and provides the best read performance without the overhead of joins or multiple round trips.

Exam trap

The trap here is assuming that any list should be referenced to avoid duplication, but MongoDB's guidance is to embed when the data is small and accessed together.

48
MCQmedium

What is the primary advantage of the Subset pattern in MongoDB data modeling?

A.It eliminates the need for indexes on the collection.
B.It keeps the working set small to improve cache performance.
C.It automatically scales the database across sharded clusters.
D.It prevents data duplication across different collections.
AnswerB

By limiting the document size to only the most essential data, you maximize the number of documents that can reside in the WiredTiger cache. A smaller working set means fewer disk I/O operations, leading to faster query response times and higher overall throughput for the database cluster's performance.

Why this answer

The Subset pattern allows you to store the most frequently accessed data in the main document while offloading the rest to another collection. This ensures that the primary document remains small and fits efficiently into the WiredTiger cache, which significantly boosts read performance for common queries. By reducing the overall document size, you also avoid the 16MB BSON limit while still keeping essential information easily accessible for the application's most frequent use cases.

Exam trap

Candidates often believe the Subset pattern is about security or data masking. They miss that its primary purpose is optimizing the working set to fit into the WiredTiger cache.

49
MCQhard

A logistics application stores shipment documents. Each shipment references a carrier by carrierId. Carriers are a small, slow-changing set (about 40 documents) that the application joins on nearly every shipment view. Which design decision is most appropriate for the carrier reference?

A.Denormalize all carrier fields into the shipment and treat the shipment copy as authoritative for carrier data.
B.Keep carrierId as a reference and resolve carrier details with $lookup, optionally caching the small carrier set in the application.
C.Embed the full carrier document in every shipment to eliminate the join entirely.
D.Store carrier details in a separate database and use a manual application-side join for every shipment read.
AnswerB

Because carriers are few and slow-changing, a reference plus $lookup is cheap, and the application can cache the entire carrier set in memory, avoiding the join on most reads. This preserves a single authoritative carrier record while keeping shipment documents small. It fits the scenario's small, stable, frequently joined data set.

Why this answer

When the referenced set is small and slow-changing, a reference is preferable to embedding because it avoids duplicating data and keeps one authoritative source. The $lookup join is inexpensive for a tiny collection, and the application can cache the whole carrier set, so shipment reads rarely pay join cost. Full denormalization or cross-database joins add complexity without benefit here.

Exam trap

The trap here is assuming that frequent joins always justify embedding, when a small, stable referenced set is exactly the case where a reference plus caching is cheaper and keeps a single source of truth.

50
MCQmedium

An e-commerce application stores products with fluctuating attributes. Which data modeling pattern is most effective for handling diverse product schemas while maintaining efficient filtering?

A.Embedding every possible attribute as a top-level field in the document.
B.Storing all product attributes as a single large JSON blob string.
C.Implementing the Attribute Pattern with an array of key-value pairs.
D.Creating a separate collection for every individual product category.
AnswerC

This pattern creates a standardized structure that allows you to index the key and value fields uniformly. It enables efficient queries across different product types without modifying the schema, ensuring that MongoDB can perform high-performance index lookups regardless of which specific attributes are stored for a given document.

Why this answer

The Attribute Pattern is the standard solution for schemas with heterogeneous fields. By transforming fields into an array of key-value pairs (k: 'color', v: 'red'), you create a uniform structure that allows for single-index coverage across all attributes. This avoids the limitations of sparse indexes and allows the application to query arbitrary attributes without needing to modify the document structure or rebuild large index sets frequently.

Exam trap

Candidates often suggest using a wildcard index for heterogeneous data, ignoring that the Attribute Pattern is specifically designed for efficient, indexable filtering on diverse, dynamic key-value pairs.

51
MCQmedium

Which of the following describes the 'Extended Reference' pattern and its primary use case?

A.Embedding all child documents to ensure total data atomicity.
B.Storing the entire child document inside the parent for faster retrieval.
C.Copying frequently accessed fields to a referenced document to minimize joins.
D.Using an array of document IDs to link parent and child documents.
AnswerC

This pattern reduces the need for expensive $lookup operations. By duplicating frequently used data fields from the 'many' side into the parent document, you can fulfill most read requests using a single query. This improves latency and reduces server load, which is a key goal in high-performance MongoDB data modeling.

Why this answer

The Extended Reference pattern involves copying frequently accessed fields from a referenced document into the current document. This eliminates the need for costly $lookup operations during read queries. It is a vital pattern for high-performance applications where latency is critical, as it trades a small amount of data redundancy for significant gains in read speed by enabling single-document retrieval for complex data views.

Exam trap

Candidates confuse the Extended Reference pattern with complete data duplication or embedding, forgetting that it only copies a select few frequently accessed fields.

52
MCQmedium

A social media platform stores each user's followers as an array of ObjectIds inside the user document. Power users can have millions of followers, and the application frequently needs to paginate through the follower list and display total counts. Which data modeling change best addresses the document growth and query performance concerns?

A.Store followers in a separate collection with one document per follower relationship, indexed on the followed user.
B.Convert the followers array into a GridFS bucket so that the follower list is chunked across multiple files.
C.Add a $slice projection to every query so only the first 100 followers are ever returned to the application.
D.Increase the document size limit by enabling the largeDocuments storage engine option on the replica set.
AnswerA

With millions of followers, embedding ObjectIds in a single user document risks exceeding the 16MB BSON limit and forces large array scans. A separate collection with an index on the followed user supports efficient pagination and counts via cursor iteration and countDocuments, keeping individual documents small and queryable.

Why this answer

When an embedded array can grow without a realistic upper bound, the document risks hitting the 16MB BSON limit and queries become costly because large arrays must be scanned. Modeling the relationship as its own collection with a supporting index keeps documents small, enables efficient pagination, and allows accurate counts using standard aggregation or count operations.

Exam trap

The trap here is assuming the 16MB BSON limit can be configured or bypassed, when it is a fixed constraint that must be designed around.

53
MCQhard

Refer to the exhibit. You have a collection where each document contains a 'tags' array. If you frequently query for documents containing specific tags, what is the best way to model and index this data?

A.Convert the tags array into a comma-separated string.
B.Use the multikey index created on the tags array.
C.Create one index per tag value to ensure performance.
D.Store tags in a separate collection and reference them using an object ID.
AnswerB

Multikey indexes are the native way MongoDB handles array indexing. When an index is created on a field that contains an array, MongoDB automatically creates entries for each element, allowing for high-speed retrieval of documents based on specific tag values without any additional configuration or complex data restructuring.

Why this answer

MongoDB automatically creates a multikey index when you index an array field. This allows the query engine to efficiently find documents that contain any of the elements in the array. This index type is highly performant for tag-based filtering because it maps every element in the array to the document location, enabling rapid searches without scanning every document in the collection.

Exam trap

Candidates frequently try to use a standard single-field index on an array, not realizing that MongoDB requires a multikey index to correctly index individual elements within an array structure.

54
MCQmedium

A team is modeling orders and customers. Orders are frequently displayed with basic customer details such as name and email, but customer profiles are large and change often. Which approach best balances read performance with avoiding duplication problems?

A.Embed the entire customer profile document inside every order.
B.Normalize customers and orders into separate collections with no duplicated fields and use $lookup for every read.
C.Embed a small subset of frequently accessed customer fields in the order while keeping the full profile in the customers collection.
D.Store only a customerId reference in each order and always perform an application-side join.
AnswerC

Embedding a small, stable subset of customer fields makes the common read self-contained and fast, while the full profile remains in one place to avoid duplicating large or volatile data. This balances read performance against duplication and update cost. It directly matches the scenario's emphasis on frequently displayed basic details.

Why this answer

A selective embedding of only the fields commonly shown with orders keeps those reads fast and self-contained, while the authoritative, larger profile stays in its own collection to avoid widespread duplication. This balances read locality with manageable update cost. The other choices either over-duplicate volatile data or force a lookup on every read.

Exam trap

The trap here is choosing between all-or-nothing embedding and referencing, when a partial embed of frequently read fields is often the right middle ground.

55
Multi-Selectmedium

Which TWO of the following scenarios are optimal for using the Embedding pattern instead of Referencing?

Select 2 answers
A.A blog post containing a list of comments that grows indefinitely.
B.A user document containing the user's primary shipping address.
C.A product catalog with millions of items linked to various categories.
D.A movie document embedding the names of its three main cast members.
E.A sensor data stream that stores millions of readings per hour.
AnswersB, D

Addresses have a strong 'contains' relationship with the user and are rarely queried independently of the user profile. Embedding them allows the application to retrieve both the user identity and their shipping information in a single read operation, significantly reducing application-side latency and database server load.

Why this answer

Embedding is preferred when data is accessed together or exhibits a 'contains' relationship, minimizing round trips. It is ideal for one-to-few relationships. Conversely, Referencing is better for large, unbounded datasets or many-to-many relationships.

Understanding these trade-offs is critical for MongoDB developers to minimize I/O and maximize the efficiency of read operations, as embedded documents are retrieved in a single read operation by the database engine.

Exam trap

Many candidates choose embedding for unbounded or many-to-many datasets, ignoring MongoDB document size limits and performance degradation risks.

56
MCQhard

A social media application stores user profiles with an embedded `friends` array containing friend IDs. The array can grow to thousands of entries. Queries frequently need to find mutual friends between two users. Which schema design strategy best optimizes this workload in MongoDB?

A.Keep the embedded `friends` array and create a multikey index on it to support mutual friend queries.
B.Store the `friends` array as a GridFS bucket to bypass the 16MB document limit and query it with aggregation.
C.Use a `$lookup` aggregation to join the user's `friends` array with another user's `friends` array at query time.
D.Model friendships as separate documents in a `friendships` collection with fields `user1` and `user2`, and create a compound index on both fields.
AnswerD

Storing each friendship as a separate document with a compound index on `user1` and `user2` allows efficient queries to find all friends of a user and to compute mutual friends via set intersection using indexed lookups. This design avoids unbounded array growth and leverages MongoDB's indexing for fast, scalable queries on large datasets.

Why this answer

Modeling friendships as separate documents with a compound index on both user fields enables efficient indexed lookups for mutual friends. This design avoids the pitfalls of unbounded arrays and leverages MongoDB's indexing to handle large-scale social graphs, ensuring queries remain performant as the dataset grows.

Exam trap

The trap here is assuming that embedding friend lists and using a multikey index is sufficient for mutual friend queries, when in fact set intersection operations on large arrays are inefficient without a dedicated relationship collection.

57
Multi-Selectmedium

You are designing a schema for a social feed where each post has a small, bounded set of reactions (like, love, wow) that are always displayed together with the post, and a potentially very large set of comments that are paginated on demand. Which TWO design choices correctly apply MongoDB schema patterns to these requirements? (Choose two.)

Select 2 answers
A.Store both reactions and comments in a single embedded array on the post to simplify rendering.
B.Store comments in a separate collection with a postId field and a compound index on { postId: 1, createdAt: -1 } for pagination.
C.Embed the reactions array directly in the post document because the set is small and always read with the post.
D.Embed all comments in the post document so pagination can be done with $slice against the embedded array.
E.Store reactions in a separate collection and aggregate them with $lookup whenever a post is rendered.
AnswersB, C

A separate comments collection keeps post documents bounded while the compound index supports efficient, ordered pagination of comments for a given post. This matches the requirement that comments be fetched on demand and can grow large. It also isolates comment writes from the post, avoiding rewrite of a large post document on every new comment.

Why this answer

Embedding is correct for bounded data read together with its parent, so reactions belong inside the post. Referencing is correct for unbounded, independently paginated data, so comments belong in their own collection with an index supporting ordered retrieval by post. The two patterns are applied to the two data sets according to their cardinality and access patterns rather than uniformly.

Exam trap

The trap here is applying a single pattern uniformly to both reactions and comments, when the deciding factor is each set's boundedness and read affinity, not a preference for embedding or referencing in general.

58
MCQmedium

What is the primary risk of using the 'Linking Pattern' improperly?

A.It causes data inconsistency.
B.It reduces read performance due to increased $lookup overhead.
C.It forces the use of fixed schema documents.
D.It prevents the use of secondary indexes on the collection.
AnswerB

Every $lookup requires the database to scan an index or collection to resolve the reference, which is significantly slower than reading an embedded document. Frequent lookups lead to high latency. This is the main performance trade-off developers make when they choose to normalize data across collections instead of embedding it.

Why this answer

The primary risk is creating too many 'joins' (lookups) in the application, which leads to slow performance. When you link data across collections, the database must perform extra work to resolve these references. If done too frequently, especially in high-traffic read paths, it creates significant latency and increases the load on the database, which directly undermines the performance benefits usually associated with MongoDB's document-based architecture.

Exam trap

Candidates often believe normalizing data completely into separate collections avoids all performance issues, failing to realize that excessive linking hurts throughput.

59
MCQhard

Refer to the exhibit. The document uses the Attribute Pattern. Why is this model superior for indexing compared to storing specs as a single sub-document `{ RAM: '16GB', CPU: 'i7' }`?

A.It reduces the document size compared to sub-documents.
B.It allows one index to cover any attribute query.
C.It enforces strict schema validation for every attribute.
D.It automatically performs joins between different products.
AnswerB

The Attribute Pattern with a multikey index on the key and value fields allows a single index to support queries on any attribute. This is superior to using sub-documents, which would require creating individual indexes for every attribute field, leading to an index explosion that degrades write performance and memory usage.

Why this answer

By using the Attribute Pattern with an array of key-value pairs, you can create a single multikey index on the 'specs.k' and 'specs.v' fields. This allows you to query any arbitrary attribute efficiently without having to create a separate index for every possible product specification. This approach provides a flexible, performant way to handle highly variable product schemas that would otherwise require hundreds of individual indexes.

Exam trap

Candidates often think the Attribute pattern is only for organization. They fail to see that the real power lies in the ability to create a single index for arbitrary queries.

Ready to test yourself?

Try a timed practice session using only Mongodb Data Modeling questions.