MongoDB · Free Practice Questions · Last reviewed May 2026
42real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
A high-traffic e-commerce application requires a schema that avoids multi-document transactions where possible to ensure maximum throughput. Which MongoDB philosophy aligns best with this requirement?
Normalize all data into separate collections to eliminate redundancy.
Implement a strict relational schema to enforce referential integrity.
Embed related data in a single document to ensure atomic updates.
Embedding data allows the application to update an entire entity in a single atomic write operation. This design pattern reduces the need for application-side joins and avoids the overhead of multi-document transactions. It is the core philosophy behind document-oriented modeling for high-performance, write-heavy workloads.
Use the GridFS specification for all user-generated content.
Refer to the exhibit. What is the primary benefit of the data structure shown?
It forces data normalization for better storage efficiency.
It allows for indexing the array to support efficient queries.
MongoDB supports multi-key indexes on array fields. By embedding tags in an array, you can create an index that allows for extremely fast lookups of all documents containing specific tags, which would be difficult and slow to perform if the tags were stored in a separate table.
It enforces referential integrity between products and tags.
It requires a multi-document transaction to update the tags.
Why does MongoDB choose the BSON format over JSON for storage?
BSON is human-readable and easier to debug than JSON.
BSON supports fewer data types than standard JSON.
BSON allows for faster traversal and efficient type identification.
The BSON format includes length headers for every element, allowing the database engine to skip over fields that are not relevant to a query. This binary structure significantly speeds up read/write performance and reduces CPU usage, as the database does not need to parse text strings to identify fields.
BSON is mandatory for communication with the shell only.
A company is migrating from a monolithic SQL database to MongoDB. They want to maintain high availability. Which feature is most critical to their success?
Horizontal sharding of the data.
Automatic failover via Replica Sets.
Replica Sets are the core mechanism for high availability. They provide data redundancy across multiple nodes and use an automated heartbeat and election process to promote a secondary node if the primary goes down, ensuring continuous operation and minimizing downtime for critical enterprise applications.
The use of the Aggregation Framework.
Strict enforcement of ACID transactions.
Which THREE of the following are core pillars of the MongoDB architecture?
Horizontal scalability through sharding.
Sharding allows MongoDB to distribute data across multiple servers, enabling the database to scale horizontally. This is a core pillar that allows the system to handle increasing data volume and throughput by adding more shards, which is critical for large-scale distributed applications and high-growth services.
Strict, static schema enforcement at the storage layer.
High availability through automatic replication.
Replication is a central architectural pillar. By maintaining multiple copies of data across different nodes, MongoDB ensures that the system can tolerate node failure without downtime. This automation of data consistency and failover is what enables the high availability required by modern mission-critical applications.
The document-based data model.
The document model is the foundation of MongoDB. It maps data to objects, allowing for intuitive development and performance optimization through embedding. This paradigm shift from rows and tables to documents is what makes MongoDB uniquely suited for a wide variety of modern development use cases.
Global dependency on a single centralized primary node.
Why does MongoDB use a 'document' as its fundamental unit of data?
It forces the use of fixed-width records for performance.
It maps naturally to application objects in code.
The document model maps directly to common object structures used in languages like JavaScript, Python, and Java. This allows developers to work with data in a format they are familiar with, reducing the overhead of translation and simplifying the process of persisting application-level objects directly into the database.
It ensures that no data can ever be duplicated.
It prevents the database from using indexes on fields.
Want more Philosophy and Features practice?
Practice this domainA DBA needs to automate log rotation on a Linux-based MongoDB server without restarting the mongod process. Which signal or command should be used to ensure the current log file is archived and a new one is started?
Send a SIGHUP signal to the mongod PID.
Run the 'rotateLogs' command from the admin database.
Send a SIGUSR1 signal to the mongod PID.
Sending a SIGUSR1 signal to the mongod process triggers the server to close the current log file, rename it with a timestamp, and open a new log file. This is the standard operational procedure for Linux environments to manage log growth without interrupting database services or requiring a full restart.
Update the 'systemLog.path' in the YAML config and run 'mongod --reconfig'.
When enabling access control on a production sharded cluster, what is the most secure method for ensuring internal authentication between cluster components such as mongos and mongod instances?
Using a shared keyfile with 600 permissions.
Enabling SCRAM-SHA-256 for all administrative users.
Configuring LDAP authorization for the __system user.
Deploying X.509 certificates for member authentication.
X.509 certificate authentication is the most robust method for internal cluster security. It uses a trusted Certificate Authority to verify the identity of each node. This prevents unauthorized nodes from joining the cluster and allows for easier certificate rotation and better compliance with modern security standards in enterprise environments.
A developer reports slow queries on a specific collection. You decide to enable the database profiler to capture only queries taking longer than 200ms. Which command configuration achieves this?
db.setProfilingLevel(1, { slowms: 200 })
Profiling level 1 instructs MongoDB to log operations that exceed the specified 'slowms' threshold. Setting this to 200 ensures that only queries slower than 200 milliseconds are recorded in the 'system.profile' collection, providing a targeted dataset for performance tuning without overwhelming the server with logging overhead.
db.setProfilingLevel(2, { slowms: 200 })
mongod --profile 200 --slowms 1
db.adminCommand({ profile: 0, slowms: 200 })
In a dedicated server environment, what is the primary reason an administrator would choose to manually decrease the 'storage.wiredTiger.engineConfig.cacheSizeGB' below the default value?
To increase the speed of BSON document compression.
To allow more memory for the filesystem cache to store indexes.
To reduce the size of the oplog on the primary node.
To accommodate other processes or containers running on the same host.
The most common reason to decrease the cache size is to ensure that other applications, such as monitoring agents, backup tools, or other database instances, have sufficient RAM to function. Without this adjustment, the mongod process might consume too much memory, leading to OOM (Out of Memory) kills.
Which THREE metrics provided by the 'mongostat' utility are most helpful for identifying a performance bottleneck caused by a high volume of concurrent write operations?
insert
The 'insert' column shows the number of insert operations per second. A high value in this column directly indicates a high write load. For a DBA, monitoring this metric alongside others helps determine if the current hardware and storage engine configuration are sufficient for the application's write throughput requirements.
getmore
qw
The 'qw' column stands for 'queue write' and represents the number of write operations waiting for a lock or storage engine availability. A consistently non-zero value in this column is a strong indicator that the server is experiencing a write bottleneck and cannot process requests fast enough.
vsize
update
The 'update' column tracks the number of update operations per second. Since updates are write operations that often involve both a read and a write to the data files and indexes, they contribute significantly to the write load. Tracking this metric is essential for understanding the overall write pressure.
A security audit requires that sensitive document data contained in query parameters must not be written to the MongoDB system logs. Which server configuration setting should be enabled?
security.authorization: enabled
security.redactClientLogData: true
Enabling 'redactClientLogData' ensures that the mongod process removes potentially sensitive information from log messages before they are written to disk. This specifically targets the data within commands, such as query filters or document fields, replacing them with placeholders to prevent data leaks through log files.
operationProfiling.mode: off
storage.wiredTiger.collectionConfig.blockCompressor: zlib
Want more Server Administration practice?
Practice this domainIn a three-node replica set with default write concern w:1, what happens to an unacknowledged write if the primary crashes immediately after receiving the write command but before replicating it to secondaries?
The write is automatically rolled back and re-applied when the old primary recovers.
The write remains in the local storage of the crashed node and becomes visible upon restart.
The write is lost, and the new primary will not contain the data.
Since the write was only acknowledged by the original primary and not replicated to the secondaries, it exists only on the crashed node. When a new primary is elected, it only contains data that was replicated. The original write is effectively discarded from the cluster's consistent view.
The replica set triggers an automatic consistency recovery to force the write to other nodes.
Refer to the exhibit. What is the primary function of the node with _id 2 in this replica set?
It stores a copy of the data and can serve read queries.
It performs automatic failover by hosting a redundant copy of the oplog.
It provides a vote during elections to help maintain a majority.
The arbiter is specifically configured to provide a voting member to the replica set without requiring the storage overhead of a data-bearing node. By participating in elections, it helps reach a majority of votes, which is required to elect a new primary during failover events in small sets.
It acts as a load balancer for incoming read requests.
When a secondary node is lagging behind the primary, what is the most significant risk to the replica set if the lag time exceeds the length of the oplog?
The secondary will trigger an automatic election to become primary.
The secondary will enter a 'RECOVERING' state and require a full resync.
When the oplog entries required for synchronization are no longer present on the primary, the secondary can no longer replicate incrementally. It transitions to a recovering state and must perform a full initial sync, which involves deleting existing data and copying the entire dataset from a healthy replica set member.
The primary will automatically reduce its write concern to accommodate the lag.
The replica set will automatically increase the oplog size to compensate.
In a replica set, what is the 'oplog' and why is it important?
A full backup of all data for disaster recovery.
A log of all read operations for performance auditing.
A capped collection that records all write operations for replication.
The oplog is indeed a capped collection that stores every modification to the data. Secondaries tail this collection continuously, pulling the operations and applying them to their own datasets. This ensures that all members of the replica set eventually reach the same state as the current primary node.
A temporary cache for incoming client write requests.
What occurs when a replica set member experiences a 'rollback'?
The node automatically merges its divergent data with the primary.
The node is permanently removed from the replica set until manually re-added.
The node discards data to align with the new primary and saves it to rollback files.
During a rollback, the node reverts its local state to the point of divergence from the current primary. The discarded operations are written to separate BSON files on the local filesystem. This process ensures the cluster returns to a consistent state while preserving the orphaned data for forensic analysis.
The node initiates a global election to force a return to its old state.
Which THREE factors can significantly impact the time taken for a new primary election?
The network latency between replica set members.
Higher network latency increases the time required for heartbeat signals to travel between nodes. If heartbeats take longer to arrive, the remaining nodes will wait longer before declaring the primary unreachable, thereby increasing the total time elapsed before a new election is triggered and a new primary is established.
The total number of documents in the largest collection.
The heartbeat timeout setting (electionTimeoutMillis).
This setting directly controls how long a secondary waits without receiving a heartbeat before it suspects the primary is offline and initiates an election. A smaller value results in faster failover, but it may also increase the risk of false positives caused by transient network spikes or congestion.
The amount of free RAM available on the primary node.
The total number of voting members in the replica set.
In a larger replica set, the consensus process requires more nodes to communicate and agree on the election outcome. Coordinating votes among a larger group naturally adds overhead to the election process, as the candidate must establish connectivity and verify status across a wider, more complex network topology.
Want more Replication practice?
Practice this domainAn application frequently queries a large collection using two fields: 'category' (equality) and 'timestamp' (range). Which index strategy provides the most efficient execution plan?
{ timestamp: 1, category: 1 }
{ category: 1 }
{ category: 1, timestamp: 1 }
Following the ESR rule, this index allows MongoDB to perform an index seek on the equality field and then efficiently traverse the range of timestamps. This significantly reduces the number of index nodes visited and documents loaded into memory, resulting in optimal query performance and resource utilization.
{ timestamp: 1 }
Which property of a field makes it a poor candidate for an index?
The field is frequently used in filter criteria.
The field contains unique identifiers like UUIDs.
The field has very low cardinality.
Low cardinality means the field has very few unique values across a large number of documents. When you search for one of these values, the index may return a huge portion of the collection, making a full collection scan more performant than using the index itself.
The field is a numeric timestamp.
When would you choose to create a Partial Index instead of a standard index?
When you need to ensure the index covers all documents.
When only a subset of data is frequently queried.
Partial indexes reduce index size by only including documents that meet a filter condition. This is highly effective when your application only queries active or specific subsets of data, leading to smaller index memory footprints and faster performance for those specific, high-frequency query patterns compared to full indexes.
To improve performance for all possible queries.
To automatically shard the collection data.
Refer to the exhibit. A developer attempts to run a query using both 'tags' and 'location' in a single $and operation. What is the expected behavior regarding index usage?
It will always use both indexes simultaneously.
It will fail because compound indexes are required.
The optimizer may choose an index intersection plan.
MongoDB's query optimizer can use index intersection to combine multiple single-field indexes to satisfy a query. It will evaluate the candidate indexes and, if the cost analysis indicates that intersection is the most efficient path, it will create a plan using both indexes for the filter.
It will ignore all indexes and scan the collection.
Refer to the exhibit. Which index is most effective for this query?
{ tags: 1 }
A multikey index on the 'tags' array allows MongoDB to map each individual element of the array to the document. This enables efficient lookup for specific values like 'red' and 'blue', making it the optimal choice for queries using the $all operator on an array field.
{ tags: 'hashed' }
{ tags: 1, _id: 1 }
No index is required for array fields.
What is the primary benefit of using a Covered Query in MongoDB?
It forces the data to stay in the WiredTiger cache.
It avoids reading the full document from disk.
When an index contains all fields required by a query, MongoDB returns the data directly from the index. This removes the need to perform a costly disk fetch to retrieve the full document, significantly reducing I/O operations and improving the overall latency and throughput of the query.
It automatically compresses all returned data.
It eliminates the need for any index on the collection.
Want more Indexing and Performance practice?
Practice this domainAn application needs to update a user's balance. Which method ensures the operation is atomic for a single document?
Read the document, update the value in application code, and save the document.
Use the $inc operator with the updateOne method.
The $inc operator performs an atomic increment on the server side. Because it is an atomic operation on a single document, it guarantees consistency even under high concurrency, making it the standard and safest method for updating numeric fields like balances in an application.
Wrap the operation in a multi-document transaction every time.
Use a global application-level mutex lock before updating.
Refer to the exhibit. An aggregation pipeline is failing with an error. How can you modify the pipeline to allow it to process the large dataset?
Set 'allowDiskUse' to true in the aggregation options.
The allowDiskUse option explicitly authorizes the aggregation pipeline to write temporary data to the _tmp directory on the server disk when the memory limit is reached. This enables processing of datasets that exceed 100MB of RAM, effectively resolving the memory limit error.
Increase the RAM allocated to the mongod process configuration.
Remove the $group stage and perform the grouping in application code.
Reduce the batch size in the cursor options.
Which connection string parameter should be used to ensure that an application only reads data from nodes that are considered current and consistent?
readPreference=secondary
readPreference=primary
Primary read preference forces all operations to go to the primary node, which is the source of truth for all writes. This guarantees the application reads the most recent, consistent data, preventing the issues associated with replication lag that occur when reading from secondary members.
readPreference=nearest
readPreference=secondaryPreferred
An application uses a time-series collection. What is the benefit of using the TTL (Time-To-Live) index feature?
It automatically creates a new collection for every time period.
It ensures data is compressed automatically to save disk space.
It allows the database to remove documents based on a timestamp field.
A TTL index enables the background removal of documents after a specified number of seconds. This automates data lifecycle management by deleting expired records based on a date-typed field, keeping the collection size manageable without requiring custom application code or manual deletion scripts.
It prevents duplicate documents from being inserted into the collection.
Which TWO of the following are essential for maintaining performance in a sharded cluster?
Selecting a shard key with high cardinality.
High cardinality ensures that data is evenly distributed across shards. A low cardinality key leads to large chunks that cannot be split, causing hotspots where specific shards receive a disproportionate amount of traffic, leading to performance degradation and uneven resource utilization across the cluster.
Hardcoding the shard primary in the application connection string.
Regularly monitoring and managing chunk splits.
As data grows, chunks must split to keep the data distributed. Without proper balancing, the cluster will suffer from uneven distribution, where some shards become overwhelmed. Monitoring split and move operations ensures that the cluster effectively manages its workload and continues to scale as data accumulates.
Disabling the balancer during all business hours.
Using the same shard key for every collection in the database.
What is the purpose of the 'write concern' in MongoDB?
To specify which primary node the write should be directed to.
To define the durability and replication requirements for a write.
Write concern allows developers to configure how many replica set members must acknowledge a write before it is considered successful. This provides fine-grained control over durability and consistency, enabling developers to trade off between write performance and the level of data safety they require.
To encrypt the data being written to the database.
To automatically compress the data before it reaches the disk.
Want more Application Administration practice?
Practice this domainA company stores IoT sensor data with a high ingestion rate. They use a monotonically increasing timestamp as the shard key. What is the most likely performance bottleneck for this cluster?
The balancer will move chunks too frequently, causing network saturation.
All new inserts will target the shard holding the highest range of values.
In ranged sharding, documents with values exceeding the current max chunk range are placed in the right-most chunk. If the shard key is a timestamp, every new record goes to this single shard. This negates the horizontal scaling benefits of sharding by creating a single point of contention.
Read operations will fail because the query router cannot locate the data.
The shard key must be unique, and timestamps often collide in high-volume systems.
Which TWO of the following are primary benefits of implementing sharding in a MongoDB environment?
Increased storage capacity by spreading data across multiple shards.
Each shard in a cluster stores a subset of the total dataset. By adding more shards, the cluster can manage significantly larger volumes of data than any individual replica set could support. This allows for near-infinite growth of the database footprint without requiring massive individual server upgrades.
Automatic failover and high availability for the entire cluster.
Improved read and write throughput through parallel processing.
Sharding allows operations to be distributed across multiple servers. When queries include the shard key, the mongos can route them to specific shards, enabling parallel processing of requests. This prevents any single machine from becoming a bottleneck for the entire application traffic, significantly boosting overall performance.
Simplified backup and recovery processes for large datasets.
Reduced latency for all queries regardless of the shard key used.
A developer chooses Hashed Sharding for a collection. What is a significant limitation of this sharding strategy compared to Ranged Sharding?
Hashed sharding does not support compound shard keys.
Hashed sharding requires the shard key field to be a numeric type.
Range-based queries on the shard key will result in broadcast operations.
In hashed sharding, documents with similar shard key values are unlikely to be stored on the same shard. Consequently, when a query specifies a range of values, the mongos cannot identify a subset of shards and must instead query every shard in the cluster to retrieve the results.
Hashed sharding cannot be used with the balancer to move chunks.
A DBA needs to ensure that chunk migrations only occur during a specific maintenance window from 02:00 to 04:00. Which configuration change is required?
Update the config.settings collection with an activeWindow document.
The balancer's schedule is controlled by modifying the 'activeWindow' field within the 'balancer' document of the 'config.settings' collection. This document accepts a start and stop time, and the balancer will only perform migrations during this interval, effectively managing the cluster's resource usage during production hours.
Use a cron job to call sh.stopBalancer() and sh.startBalancer().
Modify the mongod configuration file to include a balancing schedule.
Set a TTL index on the chunks collection to expire migrations.
An application serves users in Europe and North America. The DBA wants to ensure that European user data stays on shards located in Europe to comply with data residency laws. Which feature should be used?
Replica set tags for read preferences.
Zones and Zone Ranges.
Zones allow you to segment a sharded cluster based on shard key values. By assigning shards to zones (e.g., 'EU_Zone') and mapping specific key ranges to those zones, the balancer ensures that documents are stored only on the shards assigned to the matching zone, fulfilling data locality requirements.
Hashed sharding with a geo-spatial index.
The movePrimary command for each database.
Which THREE components are required to form a functional MongoDB sharded cluster?
One or more shards to store the data.
Shards are the physical servers or replica sets that contain the subset of the sharded data. They provide the actual storage and processing power for the dataset. In a production environment, each shard is implemented as a replica set to ensure high availability and data redundancy.
A replica set of config servers.
Config servers maintain the metadata for the cluster, including the mapping of chunks to shards. Since version 3.4, config servers must be deployed as a replica set. This metadata is essential for the mongos instances to know where to route queries and for the balancer to manage data.
One or more mongos query routers.
The mongos instances act as the entry point for application requests. They do not store data but instead use the metadata from config servers to route reads and writes to the correct shards. Applications connect to mongos just as they would to a standalone mongod or replica set.
A dedicated arbiter node for each shard.
A central MongoDB Ops Manager instance.
Want more Sharding practice?
Practice this domainWhen using the find() method, which cursor modifier is used to skip a specific number of documents before returning results?
.offset()
.limit()
.skip()
The skip() method is the correct cursor modifier to ignore a specified number of documents. It allows developers to paginate through large collections by defining the starting offset. When combined with sort() and limit(), it provides a reliable way to retrieve subsets of data for user interfaces.
.next()
Refer to the exhibit. Which documents will this query return?
All documents with 'electronics' tag sorted by price ascending.
The 5 cheapest products tagged 'electronics' that cost less than 500.
The query correctly filters by the tag and price range, sorts by price in ascending order to find the cheapest items, and restricts the output to the top five results. This accurately describes the logic executed by the MongoDB query engine based on the provided find, sort, and limit.
Any 5 products with price less than 500, regardless of tags.
The 5 most expensive products tagged 'electronics' that cost less than 500.
Which operator would you use to remove an element from an array field in a document?
$remove
$pop
$pull
$pull is the correct operator to remove all elements from an array that match a specific condition. It is efficient and atomic, allowing for precise modification of array fields without the need to read the full document state into the application, which is crucial for high-concurrency systems.
$unset
What happens when you perform a find() query on a collection where the projection includes both an inclusion and an exclusion for different fields?
The query returns an error.
MongoDB forbids mixing inclusion and exclusion in a single projection, as it creates ambiguity about which fields should be returned. The database server will explicitly return an error when it detects this configuration, forcing the developer to choose a consistent projection strategy for the requested fields.
The inclusion takes precedence over the exclusion.
The exclusion takes precedence over the inclusion.
It returns the intersection of both fields.
What is the primary function of the 'projection' parameter in the find() method?
To filter which documents are returned.
To specify which fields to return in the documents.
The projection defines the shape of the returned document by including or excluding specific fields. It is a critical performance tool because transferring unnecessary data from the database to the application increases latency and memory usage, especially when dealing with large collections or documents with many fields.
To sort the returned documents.
To limit the number of documents returned.
Refer to the exhibit. What is the effect of the 'writeConcern' parameter in this operation?
It forces the update to occur only on the primary node.
It ensures the data is replicated to most nodes before acknowledgment.
Setting w: 'majority' ensures that the write operation is committed to a majority of the voting members in the replica set. This provides a high level of data durability, ensuring that if a failover occurs, the update will persist on the newly elected primary node.
It only updates documents that have a majority of fields set.
It makes the update operation read-only.
Want more CRUD Operations practice?
Practice this domainThe C100DBA exam has 60–90 questions and must be completed in 120 minutes. The passing score is 700/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 7 domains: Philosophy and Features, Server Administration, Replication, Indexing and Performance, Application Administration, Sharding, CRUD Operations. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official MongoDB C100DBA exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.