Courseiva

C100DEV · topic practice

Data Modeling practice questions

This domain covers MongoDB schema design: embedding versus referencing, array growth, document size limits, and patterns like Subset. Questions present a relationship or workload and ask you to choose a schema, or ask what breaks when a design choice is wrong. Expect scenario-based items tied to the 16MB BSON document limit and unbounded arrays.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Data Modeling

What the exam tests

What to know about Data Modeling

You must be able to pick embedding or referencing for a given one-to-many relationship and justify it by access pattern and cardinality. The single most important thing: never let a document or embedded array grow unbounded, because the 16MB BSON limit will break writes.

Choosing embedding versus referencing based on access patterns, cardinality, and whether data is read together

Recognizing that a BSON document cannot exceed the 16MB limit and what that means for embedded arrays

Applying the Subset pattern to keep frequently accessed fields in a smaller working set

Identifying unbounded array growth as a schema risk and using the Bucket or Outlier pattern

Watch out for

Common Data Modeling exam traps

  • ▸Assuming embedding is always faster, ignoring that a document growing past 16MB causes write failures and forces schema redesign.
  • ▸Treating one-to-many as always embed or always reference, instead of deciding from cardinality and query access patterns.
  • ▸Embedding an ever-growing array such as comments or events, which bloats documents and degrades reads and writes over time.

Practice set

Data Modeling questions

20 questions · select your answer, then reveal the explanation

Which THREE factors should be considered when choosing between Embedding and Referencing in MongoDB?

Question 2easymultiple choice
Read the full Data Modeling explanation →

Which of the following describes the 'Extended Reference' pattern in MongoDB?

When should you NOT use the Embedding pattern? Select TWO.

Question 4easymultiple choice
Read the full Data Modeling explanation →

Which pattern is best suited for storing data that changes rarely but is needed in many different queries?

Which TWO of the following scenarios are best suited for the Embedding pattern rather than Referencing?

Which THREE factors should you evaluate when deciding whether to embed or reference data?

Question 7hardmultiple choice
Read the full Data Modeling explanation →

Refer to the exhibit. You have a document structure where orders are embedded in the user profile. As the number of orders grows, you notice performance issues during document updates. What is the most effective way to resolve this while maintaining read performance?

Exhibit

{"user_id": 101, "orders": [{"id": 1, "total": 50}, {"id": 2, "total": 100}]}
Question 8mediummultiple choice
Read the full Data Modeling explanation →

A financial analytics team stores daily trade records in MongoDB. Each document represents one trading day and contains a `trades` array of subdocuments with fields `symbol`, `price`, and `volume`. Queries typically filter by `symbol` and a date range, and the team needs to index efficiently. Which schema design consideration is most appropriate for this workload?

Question 9hardmultiple choice
Read the full Data Modeling explanation →

A financial application stores transactions. Each transaction document includes fields: account_id, amount, timestamp, and a nested object 'metadata' that contains varying fields depending on the transaction type (e.g., 'check_number' for checks, 'card_last4' for card payments). Queries frequently filter on account_id and timestamp, and also need to search within metadata fields. The team wants to optimize read performance and minimize index size. Which schema design approach is most appropriate?

A financial analytics application stores daily stock prices. Each document represents a single stock and contains an array of daily price entries. The array grows by one entry per trading day, and the application frequently queries the last 30 days of prices for a given stock. The number of tracked stocks is large, and the total data volume is expected to grow significantly. Which TWO schema design patterns are most appropriate to optimize this workload? (Choose two.)

Question 11mediummultiple choice
Read the full Data Modeling explanation →

An IoT platform stores sensor readings as one document per device per hour, with an embedded array of readings. Each document also stores a precomputed 'readingCount' and 'firstReading' timestamp. The application frequently needs the latest reading for a device. Which design consideration best supports this access pattern?

Question 12hardmultiple choice
Read the full Data Modeling explanation →

A financial application stores account documents. Each account references a customer by customerId, and customer documents are updated frequently (address, contact details, KYC status). Account queries almost always need the customer's current legal name and KYC status alongside the account balance. Which referencing design best meets the read and consistency requirements?

Question 13mediummultiple choice
Read the full Data Modeling explanation →

An e-commerce application stores products with fluctuating attributes. Which data modeling pattern is most effective for handling diverse product schemas while maintaining efficient filtering?

Question 14hardmultiple choice
Read the full Data Modeling explanation →

Refer to the exhibit. You are implementing the Subset Pattern for a user profile that tracks recent orders. As the number of orders per user grows indefinitely, which strategy prevents the 'unbounded array' anti-pattern?

Exhibit

{ "user": "alice", "orders": [ { "id": 1, "total": 50 }, { "id": 2, "total": 100 } ], "order_count": 2 }
Question 15mediummulti select
Read the full Data Modeling explanation →

Which TWO of the following scenarios are optimal for using the Embedding pattern instead of Referencing?

Question 16hardmultiple choice
Read the full Data Modeling explanation →

Refer to the exhibit. You have a collection where each document contains a 'tags' array. If you frequently query for documents containing specific tags, what is the best way to model and index this data?

Exhibit

db.collection.createIndex({ "tags": 1 })
Question 17mediummultiple choice
Read the full Data Modeling explanation →

You are designing a schema for a social media platform. A user has a 'profile' document, and they can have thousands of 'followers'. How should you model the follower relationship?

Question 18mediummultiple choice
Read the full Data Modeling explanation →

When designing a schema for a time-series dataset, why is it recommended to use the 'Bucket' pattern?

Question 19hardmultiple choice
Read the full Data Modeling explanation →

Refer to the exhibit. The document uses the Attribute Pattern. Why is this model superior for indexing compared to storing specs as a single sub-document `{ RAM: '16GB', CPU: 'i7' }`?

Exhibit

{
  "product": "laptop",
  "specs": [
    { "k": "RAM", "v": "16GB" },
    { "k": "CPU", "v": "i7" }
  ]
}
Question 20mediummultiple choice
Read the full Data Modeling explanation →

What happens if a document grows beyond the 16MB limit due to an 'Embedding' design choice?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Data Modeling sessions

Start a Data Modeling only practice session

Every question in these sessions is drawn from the Data Modeling domain — nothing else.

Related practice questions

Related C100DEV topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the C100DEV exam test about Data Modeling?
You must be able to pick embedding or referencing for a given one-to-many relationship and justify it by access pattern and cardinality. The single most important thing: never let a document or embedded array grow unbounded, because the 16MB BSON limit will break writes.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Data Modeling questions in a focused session?
Yes — the session launcher on this page draws every question from the Data Modeling domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other C100DEV topics?
Use the topic links above to move to related areas, or go back to the C100DEV question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the C100DEV exam covers. They are not copied from any real exam or dump site.