Courseiva

CCNA Sharding Questions

37 questions · Sharding topic · All types, answers revealed

1
MCQmedium

A collection has a 'status' field with only three possible values: 'active', 'pending', and 'closed'. Why is 'status' a poor choice for a shard key by itself?

A.The field is not a unique identifier for the documents.
B.Low cardinality prevents the creation of many chunks.
C.The field type must be an ObjectID for performance reasons.
D.The balancer cannot migrate chunks based on string values.
AnswerB

With only three possible values, MongoDB can only create a very limited number of chunks. Even in a cluster with ten shards, only three shards would ever hold data for this collection, while the others remain empty. This effectively caps the horizontal scalability of the database to three nodes.

Why this answer

A good shard key needs high cardinality to allow for a large number of chunks. If a shard key only has three possible values, the cluster can have at most three chunks for that collection. This limits the total number of shards that can participate in storing and processing the data, leading to scalability bottlenecks and imbalance.

Exam trap

Candidates often incorrectly assume that low cardinality is beneficial because it simplifies query routing, failing to realize that it prevents the cluster from splitting data into many chunks.

2
Multi-Selecthard

A financial services company runs a sharded MongoDB cluster with three shards. They need to shard a new collection that stores transaction records. The collection will be queried primarily by 'accountId' and 'transactionDate'. The team wants to minimize scatter-gather queries and ensure even data distribution. Which two actions should they take when choosing and implementing the shard key? (Choose two.)

Select 2 answers
A.Ensure the shard key has high cardinality and low frequency to avoid jumbo chunks and hot spots.
B.Use a compound shard key with 'accountId' as the first field and 'transactionDate' as the second field to support queries that include both fields.
C.Choose a shard key that is monotonically increasing, such as an auto-incrementing transaction ID, to simplify range queries.
D.Use a random shard key generated by the application to guarantee even distribution across shards.
E.Use a hashed shard key on 'accountId' to ensure even distribution, even though it prevents efficient range queries on 'transactionDate'.
AnswersA, B

High cardinality ensures many distinct shard key values, and low frequency means no single value dominates. This prevents large, unsplittable chunks (jumbo chunks) and uneven write distribution. For transaction records, 'accountId' combined with 'transactionDate' can provide high cardinality and low frequency if accounts are numerous and transactions are spread over time.

Why this answer

The compound shard key with 'accountId' first supports targeted queries on account and range queries on date, while also providing high cardinality and low frequency. This combination reduces scatter-gather and avoids hot spots. Hashed or random keys improve distribution but break query targeting, and a monotonically increasing key creates write hot spots.

Exam trap

The trap here is focusing only on even distribution and ignoring query patterns, leading to choices that cause scatter-gather.

3
MCQmedium

An organization wants to restrict customer data for European Union residents to a specific set of physical servers located in Frankfurt, complying with data residency regulations. Which MongoDB sharding feature should the DBA implement to achieve this?

A.Zone sharding, which associates specific shard key ranges with designated groups of shards to enforce data residency rules.
B.Dynamic range compression, which encrypts and confines specific database collections to localized storage volumes on individual nodes.
C.Global cluster replication, which synchronizes all database documents simultaneously across every shard in every continent.
D.Read preference tagging, which routes read operations to secondary members in specific data centers without altering write locations.
AnswerA

Zone sharding, which associates specific shard key ranges with designated groups of shards to enforce data residency rules. Zone sharding enables administrators to map document ranges to specific physical shards, ensuring data partitions remain confined to designated geographical locations.

Why this answer

Zone sharding allows administrators to associate specific shard key ranges with defined sets of shards called zones. By tagging European servers into an 'EU' zone and associating the region shard key range with that zone, MongoDB guarantees that EU customer data stays exclusively on those designated physical shards. This is critical for regulatory compliance and data sovereignty requirements.

Exam trap

Candidates confuse standard sharding with zone sharding, failing to realize that zone tags are strictly required to restrict data to specific physical servers for compliance.

4
MCQeasy

A MongoDB DBA is configuring a new sharded cluster. The application requires that queries on the sharded collection include the shard key to avoid scatter-gather. The DBA decides to use a ranged shard key on the field 'region'. After sharding, the DBA notices that queries filtering only on 'region' are performing well, but queries that also filter on 'city' are still scanning all shards. What is the most likely cause?

A.The 'city' field is not indexed, causing the query to scan all shards.
B.The shard key does not include 'city', so queries with both 'region' and 'city' cannot be targeted to a single shard.
C.The balancer has not yet distributed chunks evenly, so some shards have more data and cause the query to scan them.
D.The shard key 'region' has low cardinality, so queries cannot be targeted efficiently.
AnswerB

In a sharded cluster, only queries that include the shard key can be routed to specific shards. Since the shard key is only 'region', a query with both 'region' and 'city' still only uses 'region' for targeting. The 'city' filter is applied after routing, so the query may still target multiple shards if the region spans them, but it is not scatter-gather across all shards solely due to 'city'.

Why this answer

Query targeting in a sharded cluster depends on the shard key. If the shard key is only 'region', queries that filter on 'region' can be routed to specific shards. Adding 'city' does not change the routing because 'city' is not part of the shard key.

To target on both fields, the shard key would need to be compound, including both 'region' and 'city'.

Exam trap

The trap here is assuming that any indexed field can be used for shard targeting, when only the shard key determines routing.

5
MCQmedium

A production MongoDB 6.0 sharded cluster has a collection with a ranged shard key on the field `customerId`. Over time, the team notices that nearly all new documents have monotonically increasing `customerId` values, and one shard is receiving a disproportionate share of writes while the other shards remain relatively idle. The balancer is enabled and running. Which of the following is the most likely explanation for this imbalance?

A.The shard key is monotonically increasing, causing all new inserts to target the highest chunk on a single shard.
B.The shard key field is not indexed on the shards, forcing all writes to a single primary.
C.The config servers are overloaded and cannot process chunk migration commands.
D.The collection has too few chunks, so the balancer cannot distribute them evenly across shards.
AnswerA

A monotonically increasing shard key like an incrementing `customerId` means new documents always fall into the highest chunk, which resides on one shard until the balancer migrates it. This creates a hot shard for writes, a well-known anti-pattern for ranged shard keys with sequential values.

Why this answer

With a monotonically increasing shard key, new inserts always target the chunk with the highest key range. Because chunks are assigned to shards, that highest chunk lives on one shard at a time, making it a write hotspot. The balancer will eventually migrate chunks, but the stream of new inserts keeps hitting the newest chunk, so the imbalance persists.

Exam trap

The trap here is assuming the balancer can fully compensate for a poor shard key choice, when in fact a monotonic key will always create a write hotspot regardless of balancing.

6
MCQmedium

An administrator needs to distribute a high-throughput collection evenly across a sharded cluster using a hashed shard key. What is the primary characteristic of MongoDB's hashed sharding strategy that makes it effective for preventing write bottlenecks?

A.It computes an MD5 hash of the field value to distribute documents pseudo-randomly across all available chunks and shards, preventing single-shard write bottlenecks.
B.It compresses all chunk migration payloads using cryptographic hashing algorithms to minimize network bandwidth consumption during background balancing operations.
C.It automatically converts multi-key array fields into single scalar values to prevent indexing errors on distributed sharded collections.
D.It locks the target shard during write operations to guarantee strict serializable isolation levels across distributed cluster environments.
AnswerA

It computes an MD5 hash of the field value to distribute documents pseudo-randomly across all available chunks and shards, preventing single-shard write bottlenecks. By randomizing the distribution of incoming documents based on their hash values, even sequential identifiers like auto-incrementing integers or timestamps achieve optimal write throughput across the cluster.

Why this answer

Hashed sharding applies an MD5 hash function to the specified field value before routing documents. This transforms sequential or correlated inputs into pseudo-random hash values, ensuring that writes scatter uniformly across all available cluster shards. Understanding hashed sharding mechanisms enables administrators to design resilient write-heavy architectures that prevent the single-shard bottlenecks common with monotonic keys.

Exam trap

Candidates often incorrectly believe that hashed sharding provides faster read performance, confusing the write-distribution benefits of random hashing with query-routing optimization.

7
MCQmedium

A sharded cluster has a three-member replica set for its config servers. If two of the three config servers go offline, what is the immediate impact on the cluster's operations?

A.The cluster becomes read-only for all application data.
B.The mongos instances will crash and must be restarted.
C.Chunks cannot be split or migrated between shards.
D.All shards will automatically step down their primary members.
AnswerC

Operations that require updating the cluster metadata, such as chunk splits, migrations, or dropping collections, will fail because the config server replica set lacks a majority to commit writes. The cluster remains operational for standard data access, but its ability to balance or scale is temporarily frozen.

Why this answer

Config servers store the metadata and configuration settings for the entire sharded cluster. They are deployed as a replica set to ensure high availability. If the replica set loses its majority, it can no longer process writes to the metadata, which has significant implications for administrative tasks and the dynamic behavior of the cluster.

Exam trap

Candidates often think losing config server quorum completely stops all read and write operations on existing sharded data.

8
MCQmedium

An application executes a query on a sharded collection that does not include the shard key. What is the impact on the cluster performance?

A.The query will be rejected by the mongos as an invalid operation.
B.The query will only be sent to the primary shard of the database.
C.The mongos will perform a 'scatter-gather' operation across all shards.
D.The cluster will automatically create a new index to support the query.
AnswerC

In a scatter-gather operation, the mongos sends the query to all shards simultaneously. It then waits for all shards to respond, merges the results, and returns them to the application. This process is expensive in terms of network overhead and CPU usage across the entire cluster.

Why this answer

When a query does not include the shard key, the mongos router cannot determine which shard contains the requested data. As a result, it must broadcast the query to every shard in the cluster. This 'scatter-gather' operation increases resource consumption across all nodes and can lead to significant latency and scalability issues as the cluster grows.

Exam trap

Candidates often confuse a scatter-gather query with a targeted query, mistakenly thinking the mongos router sends the request only to the shard containing the matching range.

9
MCQhard

Refer to the exhibit. An administrator inspects chunk metadata for a sharded collection using a compound shard key: { region: 1, account_id: 1 }. What does the presence of MinKey and MaxKey signify in this specific chunk range definition?

A.They represent absolute lower and upper bounds for the account_id field within the EMEA region, ensuring all accounts for that region reside in this chunk.
B.They indicate that the chunk is corrupted and requires immediate repair using the validate command on the config server.
C.They specify encryption keys used by the WiredTiger storage engine to secure data chunks at rest on shard01.
D.They are temporary placeholder values assigned by the balancer during active chunk splits that will automatically expire after twenty-four hours.
AnswerA

They represent absolute lower and upper bounds for the account_id field within the EMEA region, ensuring all accounts for that region reside in this chunk. MinKey and MaxKey act as universal wildcards for the trailing compound fields, capturing all possible account identifiers associated with the EMEA region inside shard01.

Why this answer

MinKey and MaxKey represent lower and upper boundary markers in MongoDB that evaluate smaller or larger than any other BSON data type. In this compound chunk range, they ensure that all documents matching the 'EMEA' region regardless of their account_id value fall within this specific chunk. Understanding BSON boundary types is essential for troubleshooting custom chunk splits and zone sharding configurations.

Exam trap

Candidates often misinterpret MinKey and MaxKey as actual data values within the document, rather than understanding them as internal BSON boundary markers used for chunk range definitions.

10
MCQhard

A MongoDB 6.0 sharded cluster has a collection sharded with a hashed shard key on 'userId'. The operations team needs to run an aggregation that groups documents by 'userId' and calculates the total transaction amount. They want to know whether the aggregation can be optimized to run only on the shards that contain the relevant 'userId' values. Which statement describes the correct behavior?

A.The aggregation can target specific shards only if it includes a $sort stage on 'userId' before the $group stage.
B.The aggregation can target specific shards if the $match stage includes an equality condition on 'userId'.
C.The aggregation will always run on all shards because hashed shard keys prevent any query targeting.
D.The aggregation must be run with the allowDiskUse option to target specific shards.
AnswerB

With a hashed shard key, an equality match on the shard key field allows mongos to compute the hash and target the specific shard that owns that hashed value. If the aggregation pipeline begins with a $match on 'userId' with an equality condition, the router can direct the operation to a single shard, avoiding scatter-gather.

Why this answer

For a hashed shard key, mongos can target a specific shard when the query includes an equality condition on the shard key field. This is because the hash of the value determines the chunk and thus the shard. Other stages or options do not affect targeting.

Therefore, an aggregation with a $match on 'userId' with an equality can be optimized to run on a single shard.

Exam trap

The trap here is thinking that hashed shard keys never allow targeting, when equality queries on the hashed field do target a single shard.

11
MCQhard

When is it appropriate to use a hashed shard key instead of a ranged shard key in a MongoDB sharded collection?

A.When the application generates monotonically increasing write keys, and the primary objective is to distribute write operations evenly across all shards to prevent bottlenecks.
B.When the application executes frequent range-based find queries that require contiguous document retrieval across multiple chunks.
C.When the collection dataset size is guaranteed to remain under one gigabyte, eliminating the need for automated chunk balancing.
D.When compliance regulations require all database documents to be encrypted using deterministic hashing algorithms at rest.
AnswerA

When the application generates monotonically increasing write keys, and the primary objective is to distribute write operations evenly across all shards to prevent bottlenecks. Hashing sequential inputs randomizes their distribution, ensuring all shards share the write workload instead of overloading the maximum chunk.

Why this answer

Hashed shard keys are ideal when incoming write operations rely on monotonic or sequential values, such as auto-incrementing IDs or timestamps. Hashing prevents write bottlenecks by scattering insertions across all cluster shards, whereas ranged keys would force all writes onto a single shard.

Exam trap

Candidates often choose ranged shard keys for monotonically increasing fields, creating severe write bottlenecks on a single shard.

12
MCQmedium

A sharded cluster has a collection sharded on `{ location: 1 }`. A query is executed with the filter `{ location: { $in: ["NY", "CA"] } }`. How does MongoDB route this query?

A.The query is sent to all shards because `$in` cannot be targeted.
B.The query is routed to a random shard, and that shard forwards the request to the others as needed.
C.The query is routed to the primary shard for the database because `$in` is not a targeted operation.
D.The query is routed to the shards that own chunks covering the values "NY" and "CA".
AnswerD

When a query includes the shard key with `$in`, mongos can identify the specific chunks that contain those values and route the query only to the shards hosting those chunks. This is a targeted query, which minimizes the number of shards contacted and improves performance.

Why this answer

MongoDB's query router, mongos, can target queries that include the shard key, even with `$in`. It looks up the chunk ranges for the specified values and routes the query only to the shards that own those chunks. This is known as a targeted query and is efficient because it avoids contacting unnecessary shards.

Exam trap

The trap here is assuming that `$in` queries cannot be targeted, when in fact mongos can route them to specific shards based on the shard key values.

13
MCQhard

When designing a compound shard key in MongoDB, what is the architectural significance of placing a high-cardinality field as the prefix compared to placing a low-cardinality field first?

A.A high-cardinality prefix ensures fine-grained chunk boundaries and even data distribution, whereas a low-cardinality prefix causes massive chunks that cannot be split effectively.
B.A high-cardinality prefix allows the query router to bypass authentication checks on secondary shards, improving overall cluster query response times.
C.A low-cardinality prefix enables the WiredTiger storage engine to cache entire indexes in RAM, eliminating disk I/O bottlenecks entirely.
D.A low-cardinality prefix automatically triggers hashed shard evaluation, bypassing the need for explicit hashing functions during collection creation.
AnswerA

A high-cardinality prefix ensures fine-grained chunk boundaries and even data distribution, whereas a low-cardinality prefix causes massive chunks that cannot be split effectively. Low-cardinality prefixes limit the number of distinct values, creating oversized chunks for each unique value that the balancer cannot split further, leading to severe distribution hotspots.

Why this answer

The shard key prefix dictates how MongoDB organizes and partitions chunk ranges across the cluster. A high-cardinality prefix ensures fine-grained chunk boundaries and prevents massive, unmanageable chunks from forming around repetitive values, which often happens when low-cardinality fields like boolean flags or status codes occupy the leading position.

Exam trap

Candidates often mistakenly prioritize field names or data types over cardinality, failing to realize that low-cardinality prefixes lead to monolithic, unsplittable chunks that degrade cluster performance.

14
MCQhard

A DBA is troubleshooting a sharded cluster where the balancer is not migrating chunks even though there is an imbalance. The DBA runs `sh.getBalancerState()` and it returns `true`. Which of the following is the most likely reason the balancer is not migrating chunks?

A.The config servers are in a read-only state.
B.The collection has no indexes on the shard key, so the balancer cannot identify chunks to migrate.
C.The balancer window is configured to a time period outside of the current time.
D.The shard key is hashed, and hashed shard keys prevent the balancer from migrating chunks.
AnswerC

Even if the balancer is enabled, it only runs during the configured balancing window. If the window is set to a time outside the current time, no migrations will occur. The DBA should check `sh.getBalancerWindow()` to see the active window and adjust it if necessary.

Why this answer

The balancer only runs during the configured balancing window. Even if `sh.getBalancerState()` returns `true`, migrations will not occur outside that window. The DBA should check the balancer window settings and ensure they align with the desired migration times.

Other options would either cause broader failures or are not relevant to the balancer's operation.

Exam trap

The trap here is assuming that a `true` balancer state means the balancer is actively migrating chunks at all times, ignoring the effect of the balancing window.

15
Multi-Selecteasy

Which TWO of the following are primary benefits of implementing sharding in a MongoDB environment?

Select 2 answers
A.Increased storage capacity by spreading data across multiple shards.
B.Automatic failover and high availability for the entire cluster.
C.Improved read and write throughput through parallel processing.
D.Simplified backup and recovery processes for large datasets.
E.Reduced latency for all queries regardless of the shard key used.
AnswersA, C

Each shard in a cluster stores a subset of the total dataset. By adding more shards, the cluster can manage significantly larger volumes of data than any individual replica set could support. This allows for near-infinite growth of the database footprint without requiring massive individual server upgrades.

Why this answer

Sharding provides horizontal scaling, allowing a database to handle loads beyond the capacity of a single server. By distributing data across multiple machines, it increases both the total storage capacity and the aggregate I/O throughput. This architectural pattern is essential for large-scale applications where vertical scaling becomes cost-prohibitive or physically impossible due to hardware limitations on a single node.

Exam trap

Candidates often incorrectly select 'automatic data encryption' or 'high availability' as primary benefits of sharding, failing to distinguish between core sharding features and replica set capabilities.

16
MCQhard

A multi-document transaction is being executed across multiple shards. Which component is responsible for coordinating the transaction and ensuring the all-or-nothing property?

A.The primary node of the config server replica set.
B.The mongos router that initiated the transaction.
C.The shard that contains the first document accessed in the transaction.
D.A specialized 'Transaction Manager' node that must be configured.
AnswerB

The mongos instance serves as the coordinator for the distributed transaction. It tracks which shards are involved, manages the transaction's lifecycle, and executes the commit or abort protocol across all participants. This centralized coordination is necessary to guarantee that the transaction either commits on all shards or none.

Why this answer

Multi-document transactions in a sharded cluster are coordinated by the mongos router where the transaction was initiated. The mongos acts as the transaction coordinator, using a two-phase commit protocol to ensure consistency across the participant shards. This allows applications to maintain ACID guarantees even when data is distributed across a horizontally scaled environment.

Exam trap

Candidates often mistakenly believe that the primary shard or the config servers coordinate multi-document transactions, ignoring the central role of the mongos router in the process.

17
MCQmedium

A DBA needs to ensure that chunk migrations only occur during a specific maintenance window from 02:00 to 04:00. Which configuration change is required?

A.Update the config.settings collection with an activeWindow document.
B.Use a cron job to call sh.stopBalancer() and sh.startBalancer().
C.Modify the mongod configuration file to include a balancing schedule.
D.Set a TTL index on the chunks collection to expire migrations.
AnswerA

The balancer's schedule is controlled by modifying the 'activeWindow' field within the 'balancer' document of the 'config.settings' collection. This document accepts a start and stop time, and the balancer will only perform migrations during this interval, effectively managing the cluster's resource usage during production hours.

Why this answer

MongoDB allows administrators to define a schedule for the balancer using the settings document in the config database. By setting an activeWindow, you can restrict chunk migrations to periods of low application activity. This prevents the performance overhead of data movement from impacting the user experience during peak hours while still ensuring the cluster stays balanced.

Exam trap

Candidates often incorrectly guess that the balancer is configured via a shell command or a server-wide setting, overlooking the specific requirement to modify the config.settings collection directly.

18
Multi-Selecteasy

Which THREE components are required to form a functional MongoDB sharded cluster?

Select 3 answers
A.One or more shards to store the data.
B.A replica set of config servers.
C.One or more mongos query routers.
D.A dedicated arbiter node for each shard.
E.A central MongoDB Ops Manager instance.
AnswersA, B, C

Shards are the physical servers or replica sets that contain the subset of the sharded data. They provide the actual storage and processing power for the dataset. In a production environment, each shard is implemented as a replica set to ensure high availability and data redundancy.

Why this answer

A MongoDB sharded cluster consists of three main components: Shards, Config Servers, and Query Routers (mongos). Shards store the actual data, Config Servers store the cluster's metadata and configuration, and Query Routers act as an interface for applications, directing requests to the appropriate shards. Each plays a distinct and vital role in the cluster's operation.

Exam trap

Candidates often mistakenly list 'arbiter nodes' or 'primary nodes' as required components, confusing the architectural requirements of a sharded cluster with those of a standard replica set.

19
MCQmedium

A MongoDB 6.0 sharded cluster has a collection with a ranged shard key on the field 'customerId'. The balancer has been running normally, but the operations team notices that one shard consistently holds significantly more chunks than the others. They want to understand which internal metadata collection the mongos uses to determine chunk distribution and to verify the current chunk-to-shard mapping. Which collection should they query?

A.config.shards
B.config.collections
C.admin.system.version
D.config.chunks
AnswerD

The config.chunks collection in the config database stores the mapping of each chunk to its shard, including the min and max shard key values. Querying config.chunks on a mongos or config server replica set member reveals the exact chunk distribution, allowing the team to confirm which shard owns which ranges and to diagnose imbalance.

Why this answer

Chunk distribution is tracked in the config database, specifically in config.chunks, which maps each chunk's key range to its host shard. To verify which shard holds more chunks, an administrator queries config.chunks. The other collections store different metadata: shard list, collection definitions, or version info, none of which include per-chunk placement.

Exam trap

The trap here is assuming that config.shards contains chunk distribution details, when it only lists shard hosts and states.

20
MCQmedium

An administrator notices that the balancer in a MongoDB sharded cluster is repeatedly migrating chunks between two shards, causing performance degradation. The administrator wants to temporarily stop the balancer to investigate. Which command should be used on mongos?

A.db.runCommand({balancerStop: 1}) on the admin database
B.sh.stopBalancer()
C.sh.disableBalancer()
D.sh.setBalancerState(false)
AnswerB

The sh.stopBalancer() method is the correct command to disable the balancer on a sharded cluster. It must be run on mongos and stops the balancer from initiating new chunk migrations. This allows the administrator to investigate the issue without ongoing migrations affecting performance. The balancer can be restarted later with sh.startBalancer().

Why this answer

To temporarily stop the MongoDB balancer, an administrator should run sh.stopBalancer() on mongos. This command disables the balancer and waits for any in-progress migrations to finish, ensuring a clean stop. It is the standard helper for this purpose.

The balancer can be restarted with sh.startBalancer().

Exam trap

The trap here is using the lower-level setBalancerState(false) or an invalid helper name instead of the proper sh.stopBalancer() method that cleanly stops migrations.

21
MCQmedium

A developer chooses Hashed Sharding for a collection. What is a significant limitation of this sharding strategy compared to Ranged Sharding?

A.Hashed sharding does not support compound shard keys.
B.Hashed sharding requires the shard key field to be a numeric type.
C.Range-based queries on the shard key will result in broadcast operations.
D.Hashed sharding cannot be used with the balancer to move chunks.
AnswerC

In hashed sharding, documents with similar shard key values are unlikely to be stored on the same shard. Consequently, when a query specifies a range of values, the mongos cannot identify a subset of shards and must instead query every shard in the cluster to retrieve the results.

Why this answer

Hashed sharding ensures an even distribution of data across shards by hashing the shard key values, which is excellent for handling monotonically increasing keys. However, because the hash function essentially randomizes the placement of documents, range-based queries cannot be targeted to a single shard. This choice involves a trade-off between write distribution and query efficiency for specific access patterns.

Exam trap

Candidates often assume hashed sharding is faster for all queries. They fail to realize that scattering data randomly makes range-based scans inefficient, forcing the cluster to query every shard.

22
MCQmedium

An engineering team must shard a high-volume collection tracking global financial transactions. The chosen shard key is based on a timestamp field containing the exact millisecond of each transaction. Why will this specific shard key choice severely degrade cluster write performance over time?

A.Monotonically increasing timestamp values direct all concurrent write operations exclusively to a single chunk located on the current maximum shard, creating a severe insertion bottleneck.
B.The MongoDB balancer will immediately purge all historical chunks older than twenty-four hours to conserve disk space, resulting in accidental data loss for compliance records.
C.Queries filtering by date ranges will fail because the query router cannot project timestamps across distributed shards without a secondary hashed index on the object identifier.
D.MongoDB will automatically convert the millisecond integers into floating-point numbers, causing precision mismatch errors during chunk migration coordination phases.
AnswerA

Monotonically increasing timestamp values direct all concurrent write operations exclusively to a single chunk located on the current maximum shard, creating a severe insertion bottleneck. Because the cluster cannot distribute active inserts across multiple shards simultaneously, the designated shard experiences severe resource exhaustion while all other cluster shards remain entirely underutilized during peak ingestion periods.

Why this answer

Timestamp shard keys cause write amplification on a single shard because monotonically increasing values always route insertions to the highest chunk in the cluster. This violates key sharding principles by preventing horizontal write scalability. Recognizing insert bottlenecks helps DBAs design balanced, high-throughput architectures using hashed keys or compound prefixes to distribute incoming write operations effectively across multiple cluster shards.

Exam trap

Candidates often believe that high-precision timestamps naturally distribute data well, forgetting that monotonically increasing values send 100% of writes to a single shard chunk.

23
MCQeasy

A DBA is preparing to shard a collection in a MongoDB 6.0 cluster. The collection currently has no indexes other than the default _id index. The chosen shard key is { userId: 1 }. What must the DBA do before running shardCollection?

A.Create an index on userId, because a sharded collection requires a supporting index on the shard key.
B.Convert the shard key to a hashed key first, because ranged shard keys cannot be applied to existing collections.
C.Set the collection to unsharded status and disable the balancer for the duration of the operation.
D.Move the collection to the primary shard and ensure it is empty, because only empty collections can be sharded.
AnswerA

MongoDB requires an index that starts with the shard key fields before a collection can be sharded. If no such index exists, shardCollection fails. Creating an index on userId satisfies this requirement. The index is used to enforce chunk boundaries and support routing, so it must exist prior to the sharding operation.

Why this answer

Before sharding a collection, MongoDB requires an index whose leading fields match the shard key. Without it, the shardCollection command returns an error. Creating an index on userId satisfies this requirement and enables the cluster to manage chunk boundaries and route queries.

Emptying the collection or disabling the balancer is unnecessary and would not satisfy the actual prerequisite.

Exam trap

The trap here is overlooking the index prerequisite and assuming sharding can proceed on a collection with no supporting index for the shard key.

24
Multi-Selecthard

Which TWO statements accurately describe the behavior and management of orphaned documents in a MongoDB sharded cluster? Choose 2 answers.

Select 2 answers
A.Orphaned documents are documents that remain on a shard after a chunk has been migrated away to another node in the cluster.
B.Administrators can run the cleanupOrphaned command against shard mongod instances to remove leftover documents outside active chunk ranges.
C.Orphaned documents are automatically indexed by the query router and included in all scatter-gather query results by default.
D.The cluster balancer deletes orphaned documents simultaneously across all shards during every active migration window.
E.Orphaned documents prevent the config server from performing routine replica set elections until they are manually purged.
AnswersA, B

Orphaned documents are documents that remain on a shard after a chunk has been migrated away to another node in the cluster. During chunk migrations, source shards retain copies temporarily until cleanup routines safely remove them once migration confirmation is verified.

Why this answer

Orphaned documents are leftover data chunks remaining on a source shard after a chunk migration fails to complete cleanly or before a background cleanup process runs. Understanding how MongoDB handles orphaned documents helps DBAs troubleshoot storage utilization issues and run cleanup tasks safely without risking valid data.

Exam trap

Candidates assume that once a chunk migrates, all data is instantly and completely purged from the source shard without needing any manual or background cleanup.

25
MCQmedium

A MongoDB 6.0 sharded cluster uses a shard key of { region: 1, customerId: 1 } for a global order collection. Most queries filter on customerId only, without region. The operations team observes that nearly every read is a scatter-gather across all shards. What is the most accurate explanation for this behavior?

A.The shard key fields are in the wrong order: customerId must precede region so that queries filtering only on customerId can target a subset of shards.
B.The collection uses hashed sharding, which prevents any targeted reads regardless of which shard key fields appear in the query predicate.
C.The balancer is disabled, so chunks for the customerId values are not distributed and queries cannot be routed to a single shard.
D.The shard key is compound, and MongoDB cannot route queries unless every field in the shard key is present with an equality match.
AnswerA

MongoDB can target queries using a prefix of the shard key. With customerId as the first field, an equality query on customerId matches the prefix and routes to the owning shard, eliminating scatter-gather. With region first, customerId-only queries do not match the prefix and must be broadcast, so reordering the key addresses the observed behavior.

Why this answer

Query targeting on a sharded collection depends on whether the predicate includes a prefix of the shard key. With { region: 1, customerId: 1 }, the leading field is region, so a query filtering only on customerId matches no prefix and mongos broadcasts to all shards. Placing customerId first allows equality queries on customerId to be routed to a single shard, resolving the scatter-gather pattern.

Exam trap

The trap here is assuming that any field contained in a compound shard key enables targeted reads, when only a leading prefix of the shard key can be used for routing.

26
MCQmedium

An administrator notices that the MongoDB balancer is actively migrating chunks during peak business hours, causing noticeable application latency spikes. What is the standard administrative approach to mitigate this performance impact?

A.Configure a balancer window by updating the settings collection in the config database to restrict migrations to specific off-peak hours.
B.Permanently terminate the mongos process responsible for orchestrating chunk migrations across the cluster nodes.
C.Execute the dropDatabase command on every shard to purge unbalanced chunks instantly and reset the cluster.
D.Disable journaling across all shard replica set members to accelerate background data migration speeds.
AnswerA

Configure a balancer window by updating the settings collection in the config database to restrict migrations to specific off-peak hours. Setting a defined active time window ensures that resource-intensive chunk migrations occur exclusively during scheduled maintenance periods or low-traffic intervals.

Why this answer

Configuring a balancer window restricts chunk migrations to off-peak hours, preventing background balancing traffic from contending with production application workloads. Mastering balancer scheduling controls allows DBAs to maintain cluster health and optimal data distribution without disrupting user experience during critical business hours.

Exam trap

Candidates often think they need to completely disable the balancer during business hours, missing that configuring a maintenance window is the correct granular solution.

27
MCQmedium

During routine cluster maintenance, an administrator needs to remove an underperforming shard from a production MongoDB sharded cluster safely. What is the mandatory first step the administrator must execute?

A.Run the removeShard command on the admin database via a mongos connection to initiate the automated chunk draining process.
B.Execute the dropDatabase command on every secondary member of the shard replica set to clear local storage blocks instantly.
C.Manually shut down all mongod processes associated with the target shard before notifying the config server replica set.
D.Run the reshardCollection command to merge all orphaned documents into a single backup file on the primary config server.
AnswerA

Run the removeShard command on the admin database via a mongos connection to initiate the automated chunk draining process. The removeShard command instructs the cluster balancer to begin migrating all data chunks off the specified shard onto remaining active shards in a controlled manner.

Why this answer

Removing a shard safely begins with executing the removeShard command, which initiates the draining process. The cluster balancer automatically migrates all chunks residing on the target shard to other available shards in the cluster, ensuring no user data is lost before the node is decommissioned.

Exam trap

Candidates often assume they must manually stop the balancer or move data off the shard first. In reality, the removeShard command automates the entire migration process internally once triggered.

28
MCQmedium

An application issues a find query against a sharded collection without including the shard key in the query filter. How does the mongos query router handle this operation?

A.It performs a scatter-gather query by sending the request to all shards in the cluster, merging the results before returning them to the client.
B.It rejects the query immediately with an invalid operation exception to protect cluster performance from unindexed scans.
C.It routes the query exclusively to the primary shard of the database, bypassing all other shards in the infrastructure.
D.It broadcasts the query only to the config server primary, which evaluates the predicate against cached metadata documents.
AnswerA

It performs a scatter-gather query by sending the request to all shards in the cluster, merging the results before returning them to the client. Without the shard key in the filter, mongos lacks routing clues and must query every shard, increasing network overhead and latency.

Why this answer

When a query lacks the shard key, the mongos query router cannot determine which specific shard holds the target document. Consequently, it executes a scatter-gather query, forwarding the request to every shard in the cluster, collecting the responses, and merging them before returning the result to the client.

Exam trap

Candidates incorrectly assume that queries missing the shard key will automatically fail or target only the primary shard, missing that mongos must perform an inefficient scatter-gather operation across all shards.

29
MCQhard

Refer to the exhibit. An administrator attempts to execute the shardCollection command but receives the displayed error message. What is the root cause of this failure?

A.The collection lacks a supporting index whose key pattern matches or starts with the fields specified in the shard key definition.
B.The target database has not been enabled for sharding using the enableSharding administrative command.
C.The specified shard key contains an unsupported array field that violates BSON multi-key index constraints.
D.The config server replica set is currently undergoing an election and cannot process metadata write operations.
AnswerA

The collection lacks a supporting index whose key pattern matches or starts with the fields specified in the shard key definition. MongoDB enforces this prerequisite to ensure that every sharded collection has an underlying index capable of efficiently locating document chunk boundaries.

Why this answer

MongoDB requires an explicit index whose prefix matches the designated shard key before sharding a collection. The error indicates that no such index exists on the collection, preventing the query router from establishing the necessary routing boundaries. Creating the correct index prior to running shardCollection resolves this issue.

Exam trap

Candidates often think MongoDB automatically creates the required index when sharding a collection, leading to unexpected command failures due to missing index prefixes.

30
MCQhard

An application serves users in Europe and North America. The DBA wants to ensure that European user data stays on shards located in Europe to comply with data residency laws. Which feature should be used?

A.Replica set tags for read preferences.
B.Zones and Zone Ranges.
C.Hashed sharding with a geo-spatial index.
D.The movePrimary command for each database.
AnswerB

Zones allow you to segment a sharded cluster based on shard key values. By assigning shards to zones (e.g., 'EU_Zone') and mapping specific key ranges to those zones, the balancer ensures that documents are stored only on the shards assigned to the matching zone, fulfilling data locality requirements.

Why this answer

Zones (formerly known as tag-aware sharding) allow administrators to associate specific shard key ranges with particular shards. By tagging shards with geographical identifiers and defining zones for ranges of the shard key, the balancer will automatically move data to the appropriate physical location. This is essential for meeting regulatory requirements and reducing latency.

Exam trap

Candidates often suggest using separate clusters for different regions. While possible, it is significantly more complex than using native MongoDB zones to manage data residency within one cluster.

31
MCQmedium

An application requires high write throughput for time-series data indexed by a timestamp field. You observe that all incoming data is being routed to a single shard, creating a write hotspot. Which shard key strategy best resolves this performance bottleneck?

A.Use the timestamp field as a single-field shard key.
B.Implement a range-based shard key using a UUID.
C.Apply a hashed index to the timestamp field as the shard key.
D.Increase the number of mongos instances in the cluster.
AnswerC

Hashed sharding computes a hash of the shard key value and uses that hash to determine the target shard. This ensures that even when timestamp values are strictly increasing, the documents are distributed randomly across shards, preventing the write hotspot and allowing the application to utilize the full write capacity.

Why this answer

Using a hashed shard key distributes documents uniformly across all shards in the cluster regardless of the timestamp value. By hashing the shard key, the balancer ensures that contiguous ranges of data are not sent to the same shard. This strategy is critical in MongoDB sharding to prevent write hotspots and ensure that the disk I/O and CPU utilization are evenly balanced across the entire sharded cluster, effectively scaling write operations horizontally.

Exam trap

Candidates assume that indexing a timestamp field normally prevents hotspots, forgetting that a hashed index is necessary to randomize and distribute sequential writes.

32
MCQhard

A MongoDB sharded cluster has a collection with a hashed shard key on the field 'userId'. A query is executed that includes an equality condition on userId. How does mongos route this query?

A.It computes the hash of the userId value and routes the query to the single shard that owns the chunk containing that hash value.
B.It broadcasts the query to all shards because hashed shard keys do not support targeted queries.
C.It routes the query to the primary shard of the cluster because hashed keys always map to the primary shard for reads.
D.It uses the shard key index to identify the shard but must contact the config servers to resolve the chunk location before routing.
AnswerA

With a hashed shard key, MongoDB computes a hash of the field value and uses that hash to determine chunk placement. For an equality query on the hashed field, mongos computes the same hash and can target the exact shard that owns the corresponding chunk. This makes equality queries on the hashed field efficient and targeted to a single shard.

Why this answer

For equality queries on a hashed shard key field, mongos computes the hash of the value and uses its cached chunk metadata to route the query directly to the shard owning that hash range. This enables efficient targeted queries. Scatter-gather only occurs when the query lacks an equality condition on the hashed field.

Exam trap

The trap here is believing that hashed shard keys prevent all targeted queries, when in fact equality queries on the hashed field are efficiently targeted.

33
MCQmedium

A company stores IoT sensor data with a high ingestion rate. They use a monotonically increasing timestamp as the shard key. What is the most likely performance bottleneck for this cluster?

A.The balancer will move chunks too frequently, causing network saturation.
B.All new inserts will target the shard holding the highest range of values.
C.Read operations will fail because the query router cannot locate the data.
D.The shard key must be unique, and timestamps often collide in high-volume systems.
AnswerB

In ranged sharding, documents with values exceeding the current max chunk range are placed in the right-most chunk. If the shard key is a timestamp, every new record goes to this single shard. This negates the horizontal scaling benefits of sharding by creating a single point of contention.

Why this answer

Monotonically increasing shard keys like timestamps result in all write operations being directed to a single shard, specifically the one containing the maximum range. This creates a hot shard where the CPU and I/O capacity of one node are exhausted while others remain idle. Understanding write distribution is critical for maintaining high-throughput ingest systems in MongoDB's distributed architecture.

Exam trap

Examinees mistakenly believe that monotonically increasing shard keys distribute writes evenly because new documents are continuously added to the database.

34
MCQeasy

A database administrator is setting up a new sharded cluster and needs to enable sharding on a database named 'sales' before sharding any collections. Which command should be executed on mongos to enable sharding for this database?

A.db.runCommand({enableSharding: "sales"}) on the admin database
B.sh.addShard("sales")
C.sh.shardCollection("sales", {_id: "hashed"})
D.sh.enableSharding("sales")
AnswerD

The sh.enableSharding() method is the correct command to enable sharding on a specific database. It must be run on mongos and takes the database name as an argument. Once enabled, collections within that database can be sharded using sh.shardCollection(). This is a prerequisite step before sharding any collection in the database.

Why this answer

To enable sharding on a database in MongoDB, you must run sh.enableSharding() on mongos with the database name. This marks the database as sharding-enabled in the config servers, allowing collections within it to be sharded. The other commands either shard a collection, add a shard, or use incorrect syntax.

Exam trap

The trap here is confusing the command to enable sharding on a database with the command to shard a specific collection or add a shard to the cluster.

35
MCQhard

A financial services company shards a transactions collection using a ranged shard key on an incrementing timestamp field. Over several months, the cluster develops a single hot shard that receives almost all writes, while the other shards remain nearly idle. Which change best addresses the root cause while preserving efficient range queries on timestamp?

A.Use a compound shard key that combines a low-cardinality prefix such as a hash of the timestamp bucket with the timestamp, or apply a bucketing strategy to spread inserts.
B.Enable the balancer and lower the chunk size to 32 MB so chunks migrate more aggressively during peak hours.
C.Convert the shard key to a hashed key on timestamp, which spreads writes evenly across shards and keeps range queries efficient.
D.Add more shards to the cluster so the hot shard's load is diluted across a larger number of nodes.
AnswerA

Adding a hashed or bucketed prefix distributes inserts across many chunks while preserving the timestamp as a suffix for range queries on time. This directly targets the monotonic write hotspot, keeps the cluster balanced, and still allows efficient range scans because the timestamp remains part of the key and can be used with the prefix for pruning.

Why this answer

A monotonically increasing shard key concentrates all inserts in the highest chunk, creating a single hot shard regardless of cluster size. Introducing a hashed or bucketed prefix spreads writes across many chunks and shards, while retaining timestamp as a suffix preserves efficient range queries. This combination solves the imbalance without sacrificing the query pattern the application depends on.

Exam trap

The trap here is believing that hashed sharding on the timestamp alone is acceptable, when it actually eliminates the efficient range-query capability the application requires.

36
Multi-Selecthard

Which THREE components are mandatory architectural parts of every fully functioning MongoDB sharded cluster deployment? Choose 3 answers.

Select 3 answers
A.A config server replica set that stores all cluster metadata, chunk ranges, and routing configuration data.
B.One or more mongos query router instances that interface between client applications and the sharded cluster.
C.Multiple shard replica sets that independently store partitions of the application dataset.
D.A dedicated primary balancer coordinator node running outside the config server replica set to manage migrations.
E.An independent Apache ZooKeeper quorum cluster to coordinate distributed locking across all participating mongod nodes.
AnswersA, B, C

A config server replica set that stores all cluster metadata, chunk ranges, and routing configuration data. The config server tier is mandatory because it maintains the authoritative source of truth regarding database metadata, security settings, and precise chunk range mappings across the entire cluster.

Why this answer

A sharded cluster relies on three core functional components: config servers to maintain cluster metadata, mongos query routers to direct client requests, and shard replicas to store partitioned data safely. Knowing these architecture requirements helps administrators design resilient, highly available production clusters with proper redundancy at every layer of the infrastructure.

Exam trap

Candidates often overlook the config server replica set as a mandatory component, assuming a sharded cluster only requires mongos routers and shard nodes.

37
MCQmedium

A MongoDB DBA is managing a sharded cluster where the collection 'orders' is sharded on the field 'orderDate' using ranged sharding. The DBA needs to add a new shard to the cluster to handle increasing data volume. After adding the shard, the balancer begins migrating chunks. During this migration, the DBA runs a query that includes 'orderDate' and notices increased latency. What is the most likely reason for the increased latency?

A.The config servers are overloaded with metadata updates, causing mongos to delay query routing.
B.The query must now be routed to the new shard, which has not yet been indexed, causing a full collection scan.
C.The shard key 'orderDate' is monotonically increasing, so all new data is going to the new shard, causing a hot spot.
D.The balancer migration process consumes resources on the shards and may cause temporary performance degradation.
AnswerD

Chunk migrations involve copying data from the source shard to the destination shard and updating metadata. This process consumes CPU, memory, and I/O on both shards, which can increase latency for concurrent queries. The impact is temporary and typically resolves once migrations complete.

Why this answer

Adding a shard triggers the balancer to migrate chunks to distribute data. This migration process involves copying data and updating metadata, which consumes resources on the source and destination shards. As a result, concurrent queries may experience increased latency until migrations finish.

The other options describe unrelated issues or misconceptions about indexing and hot spots.

Exam trap

The trap here is attributing latency to the new shard's lack of indexes or config server overload, when the primary cause is resource contention from chunk migration.

Ready to test yourself?

Try a timed practice session using only Sharding questions.