You must be able to pick embedding or referencing for a given one-to-many relationship and justify it by access pattern and cardinality. The single most important thing: never let a document or embedded array grow unbounded, because the 16MB BSON limit will break writes.
Start practicing
Data Modeling — choose a session length
Free · No account required
Domain overview
This domain covers MongoDB schema design: embedding versus referencing, array growth, document size limits, and patterns like Subset. Questions present a relationship or workload and ask you to choose a schema, or ask what breaks when a design choice is wrong. Expect scenario-based items tied to the 16MB BSON document limit and unbounded arrays.
Exam objectives
Choosing embedding versus referencing based on access patterns, cardinality, and whether data is read together
Recognizing that a BSON document cannot exceed the 16MB limit and what that means for embedded arrays
Applying the Subset pattern to keep frequently accessed fields in a smaller working set
Identifying unbounded array growth as a schema risk and using the Bucket or Outlier pattern
Assuming embedding is always faster, ignoring that a document growing past 16MB causes write failures and forces schema redesign.
Treating one-to-many as always embed or always reference, instead of deciding from cardinality and query access patterns.
Embedding an ever-growing array such as comments or events, which bloats documents and degrades reads and writes over time.
Click any question to see the full explanation and answer options, or start a focused practice session above.
An e-commerce application stores products with fluctuating attributes. Which data modeling pattern is most effective for handling diverse product schemas while maintaining efficient filtering?
2Refer to the exhibit. You are implementing the Subset Pattern for a user profile that tracks recent orders. As the number of orders per user grows indefinitely, which strategy prevents the 'unbounded array' anti-pattern?
3Which TWO of the following scenarios are optimal for using the Embedding pattern instead of Referencing?
4Refer to the exhibit. You have a collection where each document contains a 'tags' array. If you frequently query for documents containing specific tags, what is the best way to model and index this data?
5You are designing a schema for a social media platform. A user has a 'profile' document, and they can have thousands of 'followers'. How should you model the follower relationship?
6When designing a schema for a time-series dataset, why is it recommended to use the 'Bucket' pattern?
7Refer to the exhibit. The document uses the Attribute Pattern. Why is this model superior for indexing compared to storing specs as a single sub-document `{ RAM: '16GB', CPU: 'i7' }`?
8What happens if a document grows beyond the 16MB limit due to an 'Embedding' design choice?
9You have a collection of 'Users' and a collection of 'Groups'. Users can belong to many groups, and groups can have many users. Which modeling approach is most scalable for this many-to-many relationship?
10Refer to the exhibit. Which MongoDB design pattern is represented here?
11When designing a schema for a blog application, you need to store comments for posts. The comments grow indefinitely. Which approach is most effective for long-term scalability?
12Refer to the exhibit. Which design pattern is used to store 'total_price' in the document?
13What is the primary advantage of the Subset pattern in MongoDB data modeling?
14When modeling with MongoDB, why should you avoid 'unbounded' growth in an array field?
15You have a document with an unbounded array of data. Which TWO strategies help mitigate the risk of exceeding the 16MB document size limit?
16When designing a schema for a one-to-many relationship, what is the primary factor in deciding between embedding and referencing?
17Which schema pattern is demonstrated in the exhibit provided?
18Which TWO of the following are benefits of using the Polymorphic Pattern in MongoDB?
19When should you use the 'Computed Pattern' in MongoDB?
20Which approach is best for handling a 'One-to-Squillions' relationship in MongoDB?
21What is the primary benefit of the 'Extended Reference Pattern'?
22What is the primary risk of using the 'Linking Pattern' improperly?
23You are designing a schema for an e-commerce platform where products have a variable number of attributes like color, size, and material. Which modeling approach provides the best balance of flexibility and query performance for filtering products by these dynamic attributes?
24Which TWO of the following scenarios are best suited for using the 'Subset Pattern' in MongoDB?
25Which of the following describes the 'Extended Reference' pattern and its primary use case?
26An e-commerce application models product reviews inside an array on the product document. As popular products accumulate hundreds of thousands of reviews, write operations and document fetch operations experience performance degradation. Which data modeling pattern best resolves this issue?
27You are designing a schema for a social application where each user document stores an array of their last 50 notifications. Notifications are frequently added, and the application only needs to display the most recent 50. Which schema design approach best supports this requirement while minimizing document growth?
28A social media application stores user posts. Each post document currently embeds a comments array. A single popular post can accumulate over 50,000 comments, causing the document to grow beyond 10 MB. The development team wants to avoid hitting the 16 MB document size limit while still being able to retrieve a post with its most recent 10 comments in a single query. Which schema design should they implement?
29A social media platform stores user profiles in a `users` collection. Each user document includes a `followers` array of user IDs. For highly popular accounts, this array has grown to over 2 million entries, causing documents to approach the 16 MB BSON limit and slowing read operations. The application frequently displays a user's follower count and the first 20 followers. Which schema design change best addresses this issue?
30A financial analytics team stores daily stock snapshots as one document per symbol per day in a collection named 'prices'. Each document contains an array of 60 subdocuments, one per minute, with fields 'minute', 'open', 'close', and 'volume'. The team now needs to compute the hourly average 'close' for a single symbol over a single day. Which aggregation approach is most appropriate for this schema?
31A social media platform stores each user's followers as an array of ObjectIds inside the user document. Power users can have millions of followers, and the application frequently needs to paginate through the follower list and display total counts. Which data modeling change best addresses the document growth and query performance concerns?
32A social platform stores user profiles in a 'users' collection and their posts in a 'posts' collection. Each post document embeds a small 'author' subdocument with 'userId', 'displayName', and 'avatarUrl'. A user changes their display name. The application must keep existing posts showing the new name. What is the most accurate statement about this Extended Reference design?
33A social media application stores user profiles with an embedded `friends` array containing friend IDs. The array can grow to thousands of entries. Queries frequently need to find mutual friends between two users. Which schema design strategy best optimizes this workload in MongoDB?
34A financial reporting system reads account documents that each contain a nested array of the last 90 daily balance snapshots. Analysts run aggregations that only need the current balance and account type, but the full snapshot array is being loaded on every read. Which schema pattern most directly reduces the working set size for these read-heavy analytics queries?
35A developer is designing a collection to store catalog items. Different item categories have different attributes: books have 'isbn' and 'pageCount', while electronics have 'wattage' and 'warrantyMonths'. All items share '_id', 'name', 'price', and 'category'. Which data modeling approach best fits this requirement?
36A library management system stores book information in a `books` collection. Each book document includes fields such as `title`, `author`, `ISBN`, and `genre`. The application frequently queries books by `genre` and `author`. Which index strategy is most appropriate to optimize these queries?
37A team is designing a schema for a product catalog where different product categories have entirely different attributes: books have ISBN and page count, electronics have wattage and warranty period, and clothing has size and material. All products must be searchable in a single query by name. Which data modeling approach best fits this requirement?
38A team is designing a schema for a MongoDB application that stores user profiles. Each profile includes a list of the user's favorite movies, which is typically fewer than 20 items and updated infrequently. Queries often retrieve the entire profile along with the favorite movies. Which data modeling approach is most appropriate?
39A logistics company uses MongoDB to store shipment data. Each shipment document includes a status field that can be one of several values: 'pending', 'in_transit', 'delivered', 'cancelled'. The application frequently queries shipments by status and also needs to generate reports that count shipments per status. The team wants to ensure efficient queries and minimal index overhead. Which schema design consideration is most important?
40An IoT platform ingests temperature readings from thousands of sensors. Each reading is currently stored as its own document with sensorId, timestamp, and value, producing millions of tiny documents per day. Queries typically retrieve readings for a sensor over a specific hour. Which schema design most improves storage efficiency and query performance for this access pattern?
41An e-commerce application stores orders in an `orders` collection. Each order document contains an array of `items`, with each item having `productId`, `quantity`, and `price`. The application needs to frequently generate reports that sum the total revenue per product across all orders. Which schema design or feature best supports this requirement efficiently?
42A mobile banking app stores each customer's account document with an embedded array of the last 20 transactions for quick display, while the full transaction history lives in a separate transactions collection. Which data modeling consideration most directly justifies this split?
43A development team is designing a schema for an e-commerce application that stores product data. Products have a set of core fields (name, price, description) that are common to all products, but different product categories have unique attributes (e.g., 'screen_size' for electronics, 'fabric' for clothing). Queries often filter on these category-specific attributes. The team wants a flexible schema that supports efficient queries on any attribute. Which TWO of the following design approaches are most appropriate? (Choose two.)
44A healthcare application stores patient records. Each patient document includes a `medications` array of subdocuments, each with `name`, `dosage`, and `frequency`. The application needs to query patients who are taking a specific medication with a specific dosage. The array is expected to grow to hundreds of entries per patient. Which schema design best supports efficient querying and management of this data?
45An IoT platform ingests sensor readings every second from thousands of devices. The team wants to minimize the number of documents and index entries while still querying by device and time range. Which schema design best matches these requirements?
46A team is deciding how to model the relationship between orders and the products each order contains. They want to optimize for the common query that retrieves an order along with the name and current price of every product in it. Which two modeling considerations best support this requirement? (Choose two.)
47A team is modeling a schema for an order-processing system that must support atomic updates across an order and its line items and must guarantee that reading an order never returns a partially updated set of line items. (Choose two.)
48A team is modeling orders and customers. Orders are frequently displayed with basic customer details such as name and email, but customer profiles are large and change often. Which approach best balances read performance with avoiding duplication problems?
49A social application stores posts with an embedded array of likes containing userId values. Some posts have millions of likes. The team wants to keep post documents small and still answer questions such as whether a specific user liked a post and how many likes a post has. Which two design changes best meet these goals? (Choose two.)
50A catalog stores products of many categories, each with different attributes: books have ISBN and author, while electronics have voltage and warrantyMonths. Queries always filter by category and then by category-specific attributes. Which schema approach best supports these queries?
51A social media platform stores each user's 'followers' as an array of ObjectIds inside the user document. Analytics shows that a small number of celebrity accounts have millions of followers, and these documents are approaching the 16MB BSON limit. Which schema pattern should a MongoDB developer apply to resolve this issue?
52You are modeling an IoT telemetry system in MongoDB. Each sensor device emits a reading every 5 seconds, and the application most often queries the last 24 hours of readings for a single device. You want to minimize the number of index entries and document reads per query. Which schema design should you choose?
53An inventory system stores each product's supplier as an embedded subdocument containing the supplier name, address, and contact email. The supplier's address changes frequently, and updates now require scanning and updating thousands of product documents. Which data modeling change best resolves this maintenance problem?
54A financial application stores account documents with an embedded array of transactions. The array is unbounded and documents have reached 8MB. Auditors require a complete, ordered history of every transaction. Which approach best preserves full history while keeping account documents within size limits?
55A retail catalog team stores each product as a document. Products share core fields such as sku, name, and price, but electronics carry warrantyMonths while apparel carries sizeChart. The application queries all products uniformly by category and price. Which modeling approach best fits this requirement?
56A product catalog must support thousands of products whose attributes differ significantly: books have ISBN and page count, electronics have wattage and warranty, and clothing has size and material. Queries frequently filter on these type-specific attributes. Which schema design best supports this requirement in MongoDB?
57A development team is designing a schema for a chat application where each conversation contains messages that grow continuously. Messages are read in chronological order, and the application must display the most recent 50 messages quickly while retaining the full history. Which TWO design decisions best support these requirements? (Choose two.)
58You are designing a schema for a social feed where each post has a small, bounded set of reactions (like, love, wow) that are always displayed together with the post, and a potentially very large set of comments that are paginated on demand. Which TWO design choices correctly apply MongoDB schema patterns to these requirements? (Choose two.)
59A logistics application stores shipment documents. Each shipment references a carrier by carrierId. Carriers are a small, slow-changing set (about 40 documents) that the application joins on nearly every shipment view. Which design decision is most appropriate for the carrier reference?
You must be able to pick embedding or referencing for a given one-to-many relationship and justify it by access pattern and cardinality. The single most important thing: never let a document or embedded array grow unbounded, because the 16MB BSON limit will break writes.
The Courseiva C100DEV question bank contains 59 questions in the Data Modeling domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data Modeling domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included