Courseiva
mediumMultiple Choice

CCNP Practice Question: A data center architect is designing a…

A data center architect is designing a virtualized environment to host critical applications. The design must maximize performance by allowing virtual machines (VMs) to directly access physical CPU cores and memory without hypervisor overhead for latency-sensitive workloads. Which hypervisor configuration should be used?

⚠ Common exam trap

Cisco often tests the misconception that hyper-threading or memory ballooning can improve performance for latency-sensitive workloads, when in fact these features are designed for resource efficiency and can introduce unpredictability or overhead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure NUMA pinning and CPU pinning for each VM to dedicated cores and memory nodes

CPU pinning and NUMA pinning allow virtual machines to directly access dedicated physical CPU cores and memory nodes, eliminating hypervisor scheduling overhead and ensuring low-latency access to local memory. This configuration is essential for latency-sensitive workloads in a virtualized data center, as it provides near-bare-metal performance by avoiding resource contention and cross-NUMA memory access penalties.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable hyper-threading and overcommit CPU resources

    Why it's wrong here

    Enabling hyper-threading and overcommitting CPU resources makes logical processors share the same physical execution units, and overcommitment forces the hypervisor to time-slice physical cores among multiple vCPUs. This introduces scheduling waits, cache contention, and non-deterministic execution timing, which are unacceptable for latency-sensitive workloads. While these techniques boost utilization for bursty or general-purpose VMs, they sacrifice the predictable, uncontended CPU access that real-time network functions require.

  • ✗

    Use a Type 2 hypervisor (e.g., VMware Workstation) for better isolation

    Why it's wrong here

    A Type 2 hypervisor such as VMware Workstation runs as an application on top of a host operating system, adding an extra scheduling and abstraction layer between the VM and hardware. This design introduces additional latency in interrupt delivery, I/O completion, and CPU scheduling compared to a Type 1 hypervisor that runs directly on the hardware. The host OS also controls resources, reducing performance isolation and making it impossible to guarantee deterministic low-latency behavior for enterprise network workloads.

  • ✓

    Configure NUMA pinning and CPU pinning for each VM to dedicated cores and memory nodes

    Why this is correct

    NUMA pinning assigns each VM to a specific NUMA node so that all memory allocations stay within that node's local memory, eliminating costly remote memory access over the interconnect. CPU pinning locks a VM's vCPUs to specific physical cores, preventing the hypervisor from migrating them across sockets or cores and removing run-queue and scheduling delays. Combined, these techniques provide deterministic access to dedicated cores and local memory, preserving cache locality and reducing latency to near-bare-metal levels, which is essential for latency-sensitive enterprise applications.

  • ✗

    Enable memory ballooning to reclaim unused memory from VMs

    Why it's wrong here

    Memory ballooning works by adding a balloon driver inside the guest that inflates to make the guest OS give up memory, allowing the hypervisor to reclaim those pages for other VMs. Under memory pressure, this can force the guest OS to swap or trim caches, causing unpredictable latency spikes as memory operations become slower. Ballooning is a tool for memory overcommitment, but it directly undermines performance guarantees because the VM's memory footprint is no longer fixed. For latency-sensitive workloads, memory must be reserved and never reclaimed dynamically.

About these practice questions

Courseiva writes every 350-401 question from scratch — 1,923 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This 350-401 practice question is part of Courseiva's free Cisco certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the 350-401 exam.