Courseiva

CCNA Core Ml Ai Knowledge Questions

7 of 82 questions · Page 2/2 · Core Ml Ai Knowledge topic · Answers revealed

76
MCQmedium

A team trains a transformer language model on a large corpus but the model achieves very low training loss while performing poorly on held-out text. Which action most directly addresses this outcome?

A.Lower the learning rate for the final training steps
B.Increase the number of training epochs on the same corpus
C.Apply dropout and weight decay to regularize the model
D.Reduce the size of the training corpus
AnswerC

A large gap between near-zero training loss and weak held-out performance is the classic signature of overfitting. Dropout randomly deactivates units during training, preventing co-adaptation, while weight decay constrains parameter magnitude. Together they reduce memorization of training text and improve generalization to unseen sequences, directly targeting the observed symptom.

Why this answer

Low training loss combined with weak held-out performance indicates overfitting, meaning the model memorizes training text instead of learning generalizable patterns. Regularization techniques such as dropout and weight decay constrain model capacity and penalize reliance on specific training examples, which directly reduces the train-validation gap and improves performance on unseen text.

Exam trap

The trap here is reacting to poor validation results by training longer or tuning the learning rate, when the low training loss already proves the model is memorizing rather than underfitting.

77
MCQeasy

Which of the following best defines 'Generalization' in machine learning?

A.The ability of a model to memorize training data perfectly.
B.The model's performance on previously unseen data.
C.The process of reducing the model's parameter count.
D.The speed at which a model converges during training.
AnswerB

Generalization specifically measures how well a model trained on a subset of data performs on new, independent data. High generalization performance is the hallmark of a successful model, indicating that it has learned the core relationships and features of the domain rather than simply memorizing training examples.

Why this answer

Generalization is the model's ability to perform accurately on new, unseen data that was not part of the training set. A model that performs perfectly on training data but fails on new data has failed to generalize. Achieving high generalization is the ultimate goal of all machine learning, as it ensures that the model provides value in real-world scenarios rather than just memorizing input-output patterns from a specific training batch.

Exam trap

Candidates often conflate generalization with model accuracy on training data or the ability to memorize large datasets, failing to recognize it specifically concerns performance on new, unseen data samples.

78
MCQmedium

A machine learning engineer is evaluating a generative LLM for a customer-facing question-answering system. The model produces fluent answers, but during testing it confidently states incorrect facts about company policies. The team wants a metric that specifically measures whether the model's output is supported by the provided source documents. Which evaluation approach is most appropriate?

A.Perplexity measured on a held-out set of company documents
B.ROUGE score computed against a large corpus of historical customer emails
C.BLEU score computed between the model's answers and reference answers
D.Faithfulness or groundedness evaluation that checks whether claims in the answer are entailed by the retrieved source passages
AnswerD

Faithfulness or groundedness metrics compare each claim in the generated answer against the provided source documents to determine whether the source supports it. This directly detects hallucinated policy details, which is the team's concern. It can be implemented with human annotation or with an entailment model, and it aligns with the requirement to measure support by source documents.

Why this answer

The team needs to know whether generated answers are supported by the source documents, which is precisely what faithfulness or groundedness evaluation measures. It checks entailment between answer claims and retrieved passages, catching confident hallucinations. Perplexity, BLEU, and ROUGE assess fluency or surface overlap with references, none of which guarantee that a stated policy fact is actually present in the source material.

Exam trap

The trap here is choosing a familiar text-generation metric like BLEU or ROUGE because it is easy to compute, when the scenario specifically requires verifying factual support against source documents.

79
MCQmedium

A data scientist is preparing a dataset of 50,000 customer support conversations to fine-tune a large language model. The conversations vary widely in length, and many exceed the model's maximum context window. The team wants to preserve conversational coherence while avoiding truncation that removes critical resolution details. Which preprocessing strategy is most appropriate?

A.Randomly sample a fixed number of tokens from each conversation and discard the rest
B.Truncate each conversation to the first N tokens that fit the context window
C.Chunk conversations into overlapping segments of at most N tokens, preserving order and overlap between chunks
D.Pad every conversation with a special token until it reaches the maximum context window
AnswerC

Chunking with overlap keeps each segment within the context window while preserving local coherence and ensuring that information spanning a boundary appears in at least one complete chunk. Fine-tuning on these segments lets the model learn from all parts of long conversations. Overlap reduces the chance that a resolution is cut off from its preceding problem description.

Why this answer

Long conversations that exceed the context window must be divided into smaller pieces while retaining meaning. Overlapping chunks preserve the order of turns and keep cross-boundary context visible, so the model can still learn from resolutions that depend on earlier messages. Random sampling and head truncation lose critical information, while padding does nothing for conversations that are already too long.

Exam trap

The trap here is thinking that any length-reduction method is acceptable, when the key requirement is preserving conversational coherence and the resolution details that often appear at the end of a dialogue.

80
MCQmedium

A team is pretraining a large language model on a cluster of NVIDIA GPUs. They observe that the model's training loss decreases steadily for the first few epochs but then suddenly spikes and eventually becomes NaN. They suspect this is due to exploding gradients. Which technique is most appropriate to address this issue?

A.Apply gradient clipping to limit the magnitude of gradients during backpropagation.
B.Reduce the batch size to decrease the variance of gradient estimates.
C.Increase the learning rate to help the model escape the spiking region.
D.Switch from Adam to stochastic gradient descent (SGD) without momentum.
AnswerA

Gradient clipping caps the norm or value of gradients before the optimizer step, preventing excessively large updates that destabilize training. In this scenario, the sudden loss spike and NaN indicate exploding gradients, so clipping directly mitigates the problem. It is a standard, low-overhead intervention that preserves the model architecture and data pipeline, making it the most appropriate immediate fix.

Why this answer

Exploding gradients cause large parameter updates that can destabilize training, leading to loss spikes and NaN values. Gradient clipping is a standard technique that caps gradient norms or values before the optimizer step, preventing such destructive updates. The other options either exacerbate the problem or do not directly target the root cause, making gradient clipping the most appropriate remedy for this scenario.

Exam trap

The trap here is assuming that any change to the optimizer or batch size will fix exploding gradients, when the direct and reliable solution is to clip gradients.

81
Multi-Selecthard

Which THREE of the following factors are critical when choosing a foundation model for an enterprise generative AI application?

Select 3 answers
A.Model licensing and data privacy policies
B.The number of hidden layers in the model
C.Task-specific performance and domain adaptability
D.Inference and training computational costs
E.The specific programming language used to build the model
AnswersA, C, D

In an enterprise environment, licensing and data privacy are paramount. Using models with restrictive licenses or those that send enterprise data to external APIs can pose significant legal and security risks. Understanding the provenance of the training data and how the model handles sensitive inputs is a mandatory compliance requirement.

Why this answer

Selecting a foundation model involves balancing technical performance, operational costs, and security requirements. Enterprise readiness requires understanding data privacy policies, licensing, and the ability to fine-tune the model for domain-specific tasks. Failing to evaluate these factors can lead to significant downstream issues, including legal risks, poor performance, or unmanageable infrastructure costs, making this selection process one of the most important steps in an AI project lifecycle.

Exam trap

Candidates often select purely performance-driven technical metrics like parameter size or raw benchmark scores while ignoring crucial enterprise operational constraints such as strict data privacy policies, model licensing terms, and total infrastructure costs.

82
MCQeasy

Which component of an NVIDIA AI stack is primarily responsible for providing a low-level API for high-performance collective communication primitives across multi-GPU nodes?

A.cuBLAS
B.NCCL
C.cuDNN
D.TensorRT
AnswerB

NCCL provides optimized collective communication routines that are aware of the underlying topology, such as NVLink and InfiniBand. It is the core library used by frameworks like PyTorch and TensorFlow to handle the synchronization of gradients during distributed training, ensuring maximum utilization of high-speed interconnects.

Why this answer

NCCL (NVIDIA Collective Communications Library) is the industry standard for multi-GPU communication. It is designed to provide high-performance primitives such as All-Reduce, All-Gather, and Broadcast, which are essential for distributed training and inference. Understanding NCCL is fundamental for scaling models across nodes, as it directly impacts the efficiency of gradient synchronization and model parallel strategies in high-performance computing environments.

Exam trap

Test-takers often confuse NCCL with high-level training frameworks or general container runtimes like CUDA, missing its specific role as the low-level library for multi-GPU communication primitives.

← PreviousPage 2 of 2 · 82 questions total

Ready to test yourself?

Try a timed practice session using only Core Ml Ai Knowledge questions.