Courseiva

CCNA Aio Ai Concepts Techniques Questions

66 questions · Aio Ai Concepts Techniques topic · All types, answers revealed

1
MCQmedium

An AI engineer is tuning a large language model for a summarization task. The output summaries are too verbose and include irrelevant details. Which technique should be applied to encourage concise outputs?

A.Provide a few-shot example with concise summaries
B.Use chain-of-thought prompting
C.Decrease the top-k value
D.Increase the temperature
AnswerA

Few-shot prompting supplies in-context demonstrations that steer the model's output distribution toward the desired style. By including concise summary exemplars in the prompt, the model infers the expected length and level of detail, directly countering the verbosity and irrelevant content described in the stem without retraining.

Why this answer

Providing a few-shot example with concise summaries (Option A) directly demonstrates the desired output format to the model, leveraging in-context learning to bias generation toward brevity and relevance. This is the most effective technique for controlling output style without altering the model's underlying parameters.

Exam trap

CompTIA often tests the misconception that adjusting sampling parameters (top-k, temperature) is the primary way to control output length, when in fact these parameters affect randomness and diversity, not the explicit length or relevance of the generated text.

How to eliminate wrong answers

Option B is wrong because chain-of-thought prompting encourages step-by-step reasoning, which typically increases verbosity and is designed for complex reasoning tasks, not for reducing output length. Option C is wrong because decreasing the top-k value restricts the sampling pool to the k most likely tokens, which can reduce randomness but does not inherently enforce conciseness or relevance; it may even produce repetitive or incomplete summaries. Option D is wrong because increasing the temperature raises the randomness of token selection, often leading to more diverse but also more verbose and irrelevant outputs, the opposite of the desired effect.

2
MCQmedium

A product team wants a system that can generate high-quality synthetic images of furniture in different room settings for an online catalog. The images must be photorealistic and vary in style. Which generative AI approach is BEST suited for this task?

A.Variational autoencoder (VAE)
B.Diffusion model
C.Generative adversarial network (GAN)
D.Recurrent neural network (RNN)
AnswerB

Diffusion models generate images by iteratively denoising random noise, producing photorealistic outputs with fine detail and controllable stylistic variation. This directly satisfies the requirement for high-quality, style-varied furniture images, unlike GANs' instability or VAEs' blurrier results.

Why this answer

Diffusion models are the best choice because they iteratively denoise random noise to produce high-quality, photorealistic images with diverse styles. Unlike GANs, they avoid mode collapse and training instability, and they generate more detailed and varied outputs than VAEs, making them ideal for furniture catalog images in different room settings.

Exam trap

CompTIA AI often tests the misconception that GANs are always the best for image generation, but the trap here is that GANs' mode collapse and training instability make diffusion models superior for high-quality, diverse outputs in production systems.

How to eliminate wrong answers

Option A is wrong because VAEs generate blurry images due to their variational lower bound objective, which smooths over fine details, making them unsuitable for photorealistic furniture images. Option C is wrong because GANs can suffer from mode collapse, where they generate limited variations (e.g., only one style of room), and training instability, reducing reliability for diverse catalog images. Option D is wrong because RNNs are designed for sequential data (e.g., text, time series) and cannot generate high-dimensional spatial images like furniture in room settings.

3
MCQeasy

Which machine learning paradigm is best suited for training a model to play a game by learning from its own actions and rewards, without labeled data?

A.Unsupervised learning
B.Semi-supervised learning
C.Reinforcement learning
D.Supervised learning
AnswerC

Reinforcement learning trains an agent through trial-and-error interaction with an environment, using reward signals rather than labelled examples. This matches the game-playing constraint exactly: the model learns optimal actions from its own experience and cumulative rewards, with no pre-labelled dataset required.

Why this answer

Reinforcement learning is designed for agents that learn by interacting with an environment and receiving rewards or penalties, with no labeled dataset required. The model improves its policy through trial and error, maximizing cumulative reward over time — exactly the paradigm for game-playing agents like AlphaGo and DQN.

Exam trap

AI0-001 often tests the distinction between reinforcement learning (reward-driven, no labels) and supervised learning (label-driven) — candidates pick supervised learning because games seem to have 'correct' moves, but RL learns from rewards, not labels.

How to eliminate wrong answers

Option A is wrong because unsupervised learning finds patterns or clusters in unlabeled data but has no reward signal or environment interaction, so it cannot learn a game-playing policy. Option B is wrong because semi-supervised learning uses a small amount of labeled data plus unlabeled data, which still requires labels and does not involve reward-based learning. Option D is wrong because supervised learning requires labeled input-output pairs, which are not available when an agent learns purely from its own actions and rewards.

4
MCQmedium

A developer is using a large language model via an API. They want the model to solve a math problem step by step. Which prompt engineering technique should they use?

A.Set temperature to 0.9
B.Chain-of-thought prompting
C.Few-shot prompting
D.Zero-shot prompting
AnswerB

Chain-of-thought prompting instructs the model to emit intermediate reasoning steps before the final answer, improving multi-step arithmetic accuracy. This directly satisfies the stem's requirement to solve a maths problem step by step rather than jumping to a result.

Why this answer

Chain-of-thought (CoT) prompting explicitly instructs the model to reason through intermediate steps before producing a final answer, which is exactly what's needed for multi-step math problems. By asking the model to 'think step by step,' it decomposes the problem into manageable reasoning stages, significantly improving accuracy on arithmetic and logic tasks.

Exam trap

AI0-001 often tests the confusion between few-shot prompting (providing examples) and chain-of-thought prompting (eliciting reasoning steps) — candidates see 'step by step' and pick few-shot because both involve guiding the model, but only CoT explicitly structures the reasoning process.

How to eliminate wrong answers

Option A is wrong because temperature controls randomness in token sampling — setting it to 0.9 increases creativity and variability, which is the opposite of what you want for deterministic math reasoning. Option C is wrong because few-shot prompting provides examples of input-output pairs but doesn't inherently force step-by-step reasoning; it's a demonstration technique, not a reasoning technique. Option D is wrong because zero-shot prompting simply asks the model to answer without examples or reasoning guidance, which performs poorly on complex math problems.

5
MCQeasy

A retail analytics team wants to group customers into segments based on purchase frequency, average order value, and recency, without having any predefined segment labels. They plan to use an algorithm that partitions customers into a fixed number of groups by minimizing within-cluster variance. Which technique should they use?

A.Linear regression
B.K-means clustering
C.Logistic regression
D.Principal component analysis
AnswerB

K-means clustering partitions data into a fixed number of groups by minimizing within-cluster variance, exactly matching the described goal. It is unsupervised, so no predefined labels are needed. Given features like purchase frequency, average order value, and recency, K-means will assign customers to the nearest centroid, producing the desired segments.

Why this answer

K-means clustering is the correct technique because it is an unsupervised algorithm that partitions observations into a predefined number of clusters by minimizing the sum of squared distances to cluster centroids. The team's goal of grouping customers without labels aligns perfectly with K-means. The other options are either supervised prediction methods or dimensionality reduction techniques that do not produce discrete customer segments.

Exam trap

The trap here is confusing dimensionality reduction with clustering, since principal component analysis is often mentioned alongside K-means in preprocessing pipelines but does not itself create segments.

6
Multi-Selecthard

A company wants to deploy an LLM-based chatbot that can handle sensitive customer information. Which THREE measures should be implemented to mitigate prompt injection attacks? (Choose 3)

Select 3 answers
A.Use a system prompt that instructs the model to ignore any instructions in the user input
B.Implement output filtering to detect and block harmful responses
C.Sanitize user inputs to remove special characters and escape sequences
D.Use a smaller model with fewer parameters
E.Set temperature to a low value
AnswersA, B, C

A system prompt establishes instruction hierarchy, telling the model to treat user input as data rather than commands. This directly mitigates injection attempts that try to override the chatbot's original behaviour, though it must be combined with other defences.

Why this answer

Option A is correct because a hardened system prompt that explicitly tells the model to treat user input as data and ignore embedded instructions is a primary defense against prompt injection, helping the model distinguish trusted developer instructions from untrusted user content. Option B is correct because output filtering adds a defense-in-depth layer that inspects the model's responses for harmful, leaked, or policy-violating content before it reaches the user, catching injections that bypass input controls. Option C is correct because sanitizing user inputs by stripping or escaping special characters, control tokens, and injection-style delimiters (e.g., markdown fences or role-play markers) reduces the attack surface for crafted prompts.

Option D is not correct because model size does not determine resistance to prompt injection; smaller models can still be manipulated and may even be more susceptible. Option E is not correct because temperature controls randomness in token sampling, not the model's adherence to injected instructions, so lowering it does not mitigate prompt injection.

Exam trap

AI0-001 often tests the misconception that model size or temperature settings affect security, when in fact prompt injection mitigation requires input/output controls and system-level instructions.

7
MCQmedium

An organization's AI system uses a decision tree model for loan approval. The compliance team requires explanations for each decision. Which property of decision trees makes them suitable for this requirement?

A.They can handle nonlinear relationships
B.They are robust to outliers
C.The decision rules are transparent and can be visualized as a tree
D.They can handle missing values
AnswerC

Decision trees expose their learned logic as a hierarchy of if-then splits on individual features, so each loan decision can be traced to explicit, human-readable conditions. This inherent interpretability satisfies the compliance requirement for per-decision explanations, unlike opaque ensembles or neural networks needing post-hoc surrogates.

Why this answer

Decision trees inherently provide interpretable decision rules by splitting data based on feature thresholds at each node. The entire model can be visualized as a tree structure, allowing compliance teams to trace the exact path and logic behind each loan approval or rejection, which directly satisfies explainability requirements.

Exam trap

CompTIA often tests the distinction between model performance properties (e.g., handling nonlinearity, robustness) and interpretability properties, leading candidates to select a technically true but irrelevant advantage instead of the one that directly satisfies the compliance requirement.

How to eliminate wrong answers

Option A is wrong because handling nonlinear relationships is a general capability of many models (e.g., neural networks, SVMs with kernels) and is not unique to decision trees, nor does it directly address the need for transparent explanations. Option B is wrong because decision trees are not inherently robust to outliers; in fact, they can be sensitive to outliers that cause splits to be skewed, and robustness is not related to explainability. Option D is wrong while decision trees can handle missing values through surrogate splits or other imputation methods, this property does not provide the transparency or traceability required for compliance explanations.

8
MCQmedium

A company wants to automatically group customer support tickets into categories (e.g., billing, technical, account) without pre-labeled data. Which machine learning approach should they use?

A.Supervised classification with logistic regression
B.Semi-supervised learning with a small labeled set
C.Unsupervised clustering using K-means
D.Reinforcement learning with a reward function
AnswerC

K-means partitions unlabelled tickets into k clusters by minimising within-cluster variance across feature vectors, directly satisfying the no-pre-labelled-data constraint. Unlike supervised classification, it discovers the billing, technical and account groupings from inherent similarity rather than predefined labels, matching the stated requirement for automatic categorisation.

Why this answer

The company has no pre-labeled data, which means supervised learning (which requires labeled examples) is not feasible. Unsupervised clustering, such as K-means, groups data points into clusters based on feature similarity without needing any labels, making it ideal for automatically discovering categories like billing, technical, or account from raw ticket text.

Exam trap

CompTIA AI often tests the distinction between supervised and unsupervised learning by presenting a scenario with 'no pre-labeled data' to trick candidates into choosing semi-supervised learning (Option B) because it sounds like a compromise, but the correct answer is always unsupervised clustering when zero labels are available.

How to eliminate wrong answers

Option A is wrong because supervised classification with logistic regression requires a pre-labeled training dataset, which the company does not have. Option B is wrong because semi-supervised learning still requires at least a small set of labeled data to guide the model, contradicting the 'without pre-labeled data' condition. Option D is wrong because reinforcement learning uses a reward function to learn a policy through trial-and-error interactions with an environment, which is not suited for static grouping of text data into categories.

9
MCQeasy

In unsupervised learning, which task involves grouping similar data points together based on feature similarities?

A.Anomaly detection
B.Classification
C.Clustering
D.Regression
AnswerC

Clustering partitions unlabelled data into groups whose members share similar feature values, maximising intra-cluster similarity. It is the unsupervised task defined by grouping similar points, distinguishing it from dimensionality reduction or density estimation, which do not produce discrete similarity-based groupings.

Why this answer

Clustering is the unsupervised learning task that groups similar data points based on feature similarities, without predefined labels. It identifies inherent structures in data by minimizing intra-cluster distances and maximizing inter-cluster distances. Common algorithms include K-means, DBSCAN, and hierarchical clustering.

This contrasts with supervised tasks like classification and regression, which require labeled data.

Exam trap

AI0-001 often tests the distinction between supervised and unsupervised learning tasks, and candidates may confuse clustering with classification because both involve grouping, but classification requires labeled data.

How to eliminate wrong answers

Option A is wrong because anomaly detection identifies rare or unusual data points that deviate from the norm, not grouping similar points together. Option B is wrong because classification is a supervised learning task that assigns predefined labels to data points, requiring labeled training data. Option D is wrong because regression is a supervised learning task that predicts continuous numerical values, not grouping similar data points.

10
Multi-Selecthard

An AI engineer is fine-tuning a transformer-based language model for a domain-specific task. They want to improve the model's factual accuracy and reduce hallucinations. Which THREE strategies should they consider? (Select THREE)

Select 3 answers
A.Increase the model's context window size beyond the training limit
B.Fine-tune the model on a curated domain-specific corpus
C.Use a higher temperature setting during generation
D.Apply chain-of-thought prompting for complex queries
E.Implement Retrieval-Augmented Generation (RAG)
AnswersB, D, E

Fine-tuning on a curated domain-specific corpus updates the model's weights toward accurate, in-domain factual patterns, directly reducing hallucination on that task. This satisfies the stem's goal of improving factual accuracy by grounding the model in verified domain content rather than relying on broad pretraining data.

Why this answer

Option B is correct because fine-tuning on a curated, high-quality domain-specific corpus injects accurate, task-relevant knowledge into the model's weights, which directly improves factual grounding and reduces the likelihood of fabricating domain facts. Option D is correct because chain-of-thought prompting encourages the model to decompose complex queries into intermediate reasoning steps, which empirically reduces errors and unsupported claims on multi-step or reasoning-heavy tasks. Option E is correct because Retrieval-Augmented Generation grounds generation in externally retrieved, up-to-date documents at inference time, so the model conditions its output on verifiable evidence rather than relying solely on parametric memory, which is the standard technique for reducing hallucinations.

Option A is not appropriate because simply increasing the context window beyond the training limit is not a supported or effective strategy—models cannot attend beyond their trained positional range without architectural changes or extrapolation techniques, and a larger window alone does not improve factual accuracy. Option C is incorrect because raising the temperature increases sampling randomness and diversity, which typically increases hallucination rather than reducing it; lower temperature is preferred for factual accuracy.

Exam trap

The CompTIA AI+ exam often tests the misconception that increasing randomness (higher temperature) or extending context windows beyond training limits can improve factual accuracy, when in fact these techniques degrade reliability.

11
MCQhard

A team is training a recurrent neural network (RNN) with LSTM units to predict stock prices. The validation loss is significantly higher than the training loss. Which action is MOST likely to reduce the gap?

A.Increase the number of LSTM units
B.Increase the number of training epochs
C.Reduce the sequence length
D.Increase the dropout rate in LSTM layers
AnswerD

Dropout randomly deactivates LSTM units during training, preventing the network from relying on particular hidden-state pathways and reducing variance. The validation loss exceeding training loss indicates overfitting, so increasing dropout regularises the recurrent layers and narrows that generalisation gap.

Why this answer

A significant gap between training and validation loss indicates overfitting, where the model memorizes training data but fails to generalize. Increasing dropout in LSTM layers is a regularization technique that randomly deactivates neurons during training, forcing the network to learn more robust features and reducing overfitting. This directly addresses the gap by improving validation performance.

Exam trap

AI0-001 often tests the diagnosis of overfitting versus underfitting; candidates may pick increasing epochs or units thinking more training will help, but those actions worsen overfitting when the validation loss is already higher than training loss.

How to eliminate wrong answers

Option A is wrong because increasing the number of LSTM units adds model capacity, which would likely worsen overfitting and increase the gap further. Option B is wrong because training for more epochs allows the model to memorize the training data even more, exacerbating overfitting. Option C is wrong because reducing sequence length may lose temporal dependencies and does not directly address overfitting; it could even hurt performance if the shortened sequences lack predictive information.

12
Multi-Selectmedium

A data scientist is evaluating a binary classifier for a medical diagnosis task. The dataset is imbalanced with 5% positive cases. Which THREE metrics should the data scientist consider for a comprehensive evaluation?

Select 3 answers
A.Precision
B.F1 score
C.Accuracy
D.Perplexity
E.Recall
AnswersA, B, E

Precision measures the proportion of predicted positives that are truly positive, exposing how many false alarms the classifier raises. At a 5% positive rate, this matters because accuracy alone would look high while missing poor positive-class performance.

Why this answer

Precision (A) is correct because it measures the proportion of predicted positives that are truly positive, which is critical in an imbalanced medical diagnosis task where false positives carry real cost. Recall (E) is correct because it measures the proportion of actual positives that are correctly identified, ensuring the classifier does not miss the rare 5% positive cases. F1 score (B) is correct because it is the harmonic mean of precision and recall, providing a single balanced metric that is far more informative than accuracy on skewed class distributions.

Accuracy (C) is not appropriate here because a trivial model predicting all negatives would achieve 95% accuracy while detecting zero positive cases. Perplexity (D) does not belong because it is a language-model metric measuring how well a probability distribution predicts a sample, not a classification performance measure.

Exam trap

The trap here is that candidates often default to accuracy as a universal metric, but CompTIA AI tests the understanding that accuracy is unreliable for imbalanced datasets, and that metrics like precision, recall, and F1 score are required for a comprehensive evaluation.

13
MCQeasy

An AI practitioner needs to measure the performance of a binary classification model for disease detection, where the cost of false negatives is very high. Which metric should be prioritized?

A.Recall
B.Precision
C.F1-score
D.Accuracy
AnswerA

Recall measures the proportion of actual positives correctly identified, directly minimising false negatives. In disease detection, missing a genuine case is costly, so maximising recall ensures fewer missed diagnoses, even at the expense of more false positives.

Why this answer

Recall (true positive rate) minimises false negatives, which is critical when missing a positive case is dangerous.

14
MCQmedium

A machine learning engineer is training a neural network for image classification. The training loss decreases slowly and the model accuracy improves only marginally each epoch. Which hyperparameter adjustment is MOST likely to accelerate convergence?

A.Add more hidden layers
B.Increase the batch size
C.Increase the learning rate
D.Decrease the number of epochs
AnswerC

Raising the learning rate increases the step size taken along the loss gradient, so each epoch moves weights further and convergence accelerates. The stem's slow loss decrease and marginal accuracy gains indicate steps that are too small, making a higher rate the direct fix.

Why this answer

The training loss decreasing slowly and accuracy improving marginally each epoch indicates that the learning rate is too small, causing the optimizer to take very small steps toward the minimum of the loss function. Increasing the learning rate allows the optimizer to take larger steps per update, which accelerates convergence. Option C is correct because adjusting the learning rate directly addresses the step size in gradient descent.

Exam trap

CompTIA AI often tests the misconception that adding more layers or increasing batch size always improves training speed, when in fact the learning rate is the primary hyperparameter controlling convergence rate.

How to eliminate wrong answers

Option A is wrong because adding more hidden layers increases model complexity and can lead to slower convergence or overfitting, not faster convergence. Option B is wrong because increasing the batch size reduces the variance of gradient estimates but does not directly speed up convergence; it can actually slow down training due to fewer weight updates per epoch. Option D is wrong because decreasing the number of epochs reduces training time but does not accelerate convergence per epoch; it may stop training before the model has converged.

15
MCQmedium

A machine learning engineer is training a logistic regression model and notices that the loss is decreasing very slowly. The learning rate is set to 0.001. What is the MOST likely cause and appropriate fix?

A.The learning rate is too low; increase it to 0.01
B.The learning rate is too high; decrease it to 0.0001
C.The model is overfitting; add L2 regularisation
D.The batch size is too large; reduce it
AnswerA

With a learning rate of 0.001, each gradient step barely shifts the weights, so loss falls slowly. Raising it to 0.01 increases the step size, accelerating convergence while remaining stable for logistic regression on typical scaled data.

Why this answer

A learning rate of 0.001 is very low for many logistic regression implementations, causing the gradient descent algorithm to take extremely small steps toward the minimum of the loss function. This results in a slow decrease in loss because each weight update is minimal. Increasing the learning rate to 0.01 allows larger steps per iteration, accelerating convergence without typically causing divergence in well-scaled data.

Exam trap

A common misconception is that a slow decrease in loss always indicates a learning rate that is too high, when in fact a very low learning rate is the typical cause for slow convergence.

How to eliminate wrong answers

Option B is wrong because a learning rate that is too high would cause the loss to oscillate or diverge, not decrease slowly; decreasing it further would worsen the slow convergence. Option C is wrong because overfitting manifests as low training loss but high validation loss, not as a slow decrease in training loss; L2 regularization addresses overfitting, not convergence speed. Option D is wrong because a large batch size can slow training in terms of wall-clock time per epoch but does not inherently cause the loss to decrease slowly per iteration; it actually provides more stable gradient estimates.

16
MCQmedium

A data scientist is selecting a model for a binary classification task where interpretability is critical because of regulatory requirements. The dataset has 20 features and 10,000 samples. Which model is MOST appropriate?

A.Neural network (MLP)
B.Decision tree
C.Gradient boosting machine
D.Random forest classifier
AnswerB

A decision tree produces human-readable if-then splits, directly satisfying the regulatory interpretability constraint. With 20 features and 10,000 samples it trains reliably, unlike neural networks or ensembles whose opaque internal weights would fail audit requirements.

Why this answer

A decision tree is inherently interpretable: its if-then-else splits can be visualized and explained to regulators, auditors, or customers. With only 20 features and 10,000 samples, a single tree is also computationally adequate and unlikely to overfit catastrophically if pruned. This makes it the best fit when interpretability is a hard requirement.

Exam trap

AI0-001 often tests the interpretability-vs-accuracy tradeoff — the trap is choosing a high-accuracy ensemble (random forest, GBM) when the question explicitly prioritizes regulatory explainability.

How to eliminate wrong answers

Option A is wrong because neural networks are black-box models whose internal weights are not human-interpretable, violating regulatory explainability requirements. Option C is wrong because gradient boosting machines, while powerful, are ensembles of trees whose combined predictions are difficult to explain without post-hoc tools like SHAP. Option D is wrong because random forests average many trees, sacrificing the single-tree interpretability that regulators typically demand.

17
MCQmedium

A developer is using a pre-trained BERT model for a question-answering system. They want to ensure the model can handle out-of-vocabulary words. Which component of the BERT architecture is responsible for this?

A.Positional encoding
B.Feed-forward layers
C.WordPiece tokenisation
D.Attention mechanism
AnswerC

WordPiece tokenisation splits unknown or rare words into frequently occurring subword units drawn from a fixed vocabulary, so the model represents out-of-vocabulary words as sequences of known subwords rather than a single unknown token. This is the component handling OOV input.

Why this answer

WordPiece tokenisation is the component of BERT that handles out-of-vocabulary (OOV) words by breaking them into subword units (e.g., 'playing' → 'play' + '##ing'). This allows the model to represent any word, even unseen ones, as a sequence of known subword tokens, ensuring no word is truly out of vocabulary.

Exam trap

The trap here is that candidates often associate 'handling unknown words' with the attention mechanism or positional encoding, but the CompTIA exam specifically tests the understanding that tokenisation—not the model's internal layers—is what makes BERT robust to OOV words.

How to eliminate wrong answers

Option A is wrong because positional encoding adds information about the position of tokens in a sequence, not about handling unknown words. Option B is wrong because feed-forward layers apply non-linear transformations to the attention output and do not address tokenisation or vocabulary coverage. Option D is wrong because the attention mechanism computes relationships between tokens but relies on the tokeniser to first convert input text into known subword pieces; it cannot handle OOV words on its own.

18
Multi-Selecthard

A retail bank is building a churn prediction model on 12 months of customer data. The data engineering team realizes that some features, such as total transactions in the last 90 days, are recorded at the moment the extraction job runs rather than at the moment each customer's churn label was determined. The model shows suspiciously high validation accuracy. Which TWO practices should the team adopt to obtain a trustworthy estimate of model performance? (Choose two.)

Select 2 answers
A.Standardize all numeric features using the mean and standard deviation computed over the full dataset
B.Evaluate the model with a time-based split that trains on earlier periods and validates on later periods
C.Construct features using only information that was available before each customer's label observation date
D.Increase the number of trees in the gradient boosting ensemble until validation accuracy stops improving
E.Apply SMOTE to oversample the minority churn class before splitting the data into train and test sets
AnswersB, C

A temporal split respects the chronological order of events, so the validation set consists of customers whose outcomes occur after the training period. This mimics the real deployment setting where the model predicts future churn from past behavior. Combined with point-in-time features, it exposes whether the model truly generalizes forward in time instead of exploiting patterns that only exist within a randomly shuffled dataset.

Why this answer

The inflated accuracy comes from target leakage: features were captured after the label moment, so they encode information about the outcome. Rebuilding features with point-in-time correctness removes that future knowledge, and a time-based train/validation split mirrors how the model will be used on future customers. Model tuning, resampling before splitting, and global scaling leave the leakage intact or introduce new leakage, so they cannot yield a trustworthy performance estimate.

Exam trap

The trap here is treating high validation accuracy as evidence of a good model when the real cause is feature values that were recorded after the outcome being predicted.

19
MCQmedium

A developer is building a natural language processing system to classify customer reviews as positive, neutral, or negative. They have 50,000 labeled reviews. Which model architecture is MOST appropriate for this task?

A.Use a convolutional neural network (CNN) on raw text
B.Train a recurrent neural network (RNN) from scratch
C.Fine-tune a pre-trained BERT model
D.Word2vec embeddings followed by logistic regression
AnswerC

Fine-tuning a pre-trained BERT model leverages transformer self-attention and language representations learned from vast corpora, then adapts them to three-class review sentiment using the 50,000 labelled examples, yielding strong accuracy where training a model from scratch would underperform.

Why this answer

Fine-tuning a pre-trained BERT model is most appropriate because BERT is a transformer-based model pre-trained on a large corpus and can be fine-tuned on the 50,000 labeled reviews to achieve high accuracy with relatively little data. It captures bidirectional context, which is crucial for sentiment classification, and avoids the need for training from scratch.

Exam trap

A common mistake is to assume that training from scratch or using simpler models like logistic regression is sufficient, but pre-trained transformers like BERT are the standard for achieving high accuracy with limited labeled data.

How to eliminate wrong answers

Option A is wrong because using a CNN on raw text without embeddings or pre-processing ignores the sequential and contextual nature of language, leading to poor performance on sentiment classification. Option B is wrong because training an RNN from scratch on only 50,000 samples is prone to overfitting and underperformance compared to leveraging a pre-trained model like BERT. Option D is wrong because Word2vec embeddings followed by logistic regression provides only shallow, bag-of-words-like features and cannot capture complex contextual relationships needed for nuanced sentiment analysis.

20
MCQeasy

A data scientist is building a model to predict whether a credit card transaction is fraudulent, using labeled historical data. Which machine learning paradigm is being used?

A.Reinforcement learning
B.Unsupervised learning
C.Supervised learning
D.Self-supervised learning
AnswerC

Supervised learning trains on labelled historical data, mapping inputs to known outputs. Fraud detection uses labelled transactions marked fraudulent or legitimate, so the model learns the mapping between transaction features and the fraud label, matching this paradigm exactly.

Why this answer

Supervised learning uses labeled historical data to train a model that predicts a target variable. Here, the credit card transactions are labeled as fraudulent or not, so the model learns from these labels to classify new transactions.

Exam trap

The trap is confusing supervised learning with semi-supervised or self-supervised learning; candidates may pick self-supervised because it also uses labels, but self-supervised generates labels from the data, whereas here labels are given.

How to eliminate wrong answers

Option A is wrong because reinforcement learning involves an agent learning from rewards and penalties through interaction with an environment, not from labeled historical data. Option B is wrong because unsupervised learning deals with unlabeled data, finding patterns or clusters without predefined labels. Option D is wrong because self-supervised learning generates labels from the data itself (e.g., predicting masked words), but here the labels are explicitly provided.

21
MCQhard

A generative AI model produces images from text prompts. The outputs are often blurry and lack fine details. Which model type is MOST likely being used, and which improvement would best address this issue?

A.Variational Autoencoder (VAE); switch to a diffusion model
B.Variational Autoencoder (VAE); switch to a Generative Adversarial Network (GAN)
C.Generative Adversarial Network (GAN); increase the discriminator's capacity
D.Diffusion model; use a larger batch size during training
AnswerA

Variational autoencoders optimise a variational lower bound with a Gaussian latent prior, which averages reconstructions and yields inherently blurry, low-detail images. Diffusion models instead learn to denoise iteratively, capturing high-frequency detail, so switching directly satisfies the stem's demand for sharper outputs.

Why this answer

Variational Autoencoders (VAEs) are known for producing blurry outputs because their loss function (ELBO) encourages pixel-wise averaging, which smooths out fine details. Diffusion models, by contrast, iteratively denoise a random field, learning to reconstruct high-frequency details through a multi-step reverse process, directly addressing the blurriness issue.

Exam trap

The AI0-001 exam often tests the misconception that GANs are always the best for sharp images, but the trap here is that the question specifically describes blurry outputs—a hallmark of VAEs—and the best modern improvement is a diffusion model, not a GAN.

How to eliminate wrong answers

Option B is wrong because switching from a VAE to a GAN would improve sharpness but GANs are prone to mode collapse and training instability, making diffusion models a more robust and state-of-the-art choice for fine detail generation. Option C is wrong because GANs already produce sharp images; increasing discriminator capacity would not fix blurriness (which is a VAE characteristic) and could worsen training instability. Option D is wrong because diffusion models do not inherently produce blurry outputs; using a larger batch size during training improves gradient stability but does not address a blurriness problem that is not characteristic of diffusion models.

22
MCQmedium

A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?

A.Use a larger foundation model with a longer context window and paste all documents into each prompt
B.Use Retrieval-Augmented Generation (RAG) with the policy documents indexed in a vector store
C.Fine-tune a base LLM on the policy documents monthly
D.Train a custom model from scratch on the policy documents each month
AnswerB

RAG retrieves relevant passages from the vector store at query time and supplies them to the model as context, so monthly document updates only require re-indexing, not retraining. This satisfies the constraint that the team cannot afford to retrain a model each time.

Why this answer

Retrieval-Augmented Generation (RAG) is the most appropriate approach because it allows the chatbot to answer questions based on the latest policy documents without retraining the model. By indexing the documents in a vector store and retrieving relevant chunks at query time, RAG provides up-to-date, contextually accurate answers while keeping the underlying LLM static, which avoids the cost and complexity of monthly retraining.

Exam trap

CompTIA often tests the misconception that a larger context window or fine-tuning is the only way to handle dynamic data, when in fact RAG is the scalable, cost-effective solution for frequently updated knowledge bases without retraining.

How to eliminate wrong answers

Option A is wrong because pasting all policy documents into each prompt would quickly exceed the model's context window (even with larger models, context windows are finite and costly), leading to truncated inputs, degraded performance, and high token costs. Option C is wrong because fine-tuning a base LLM monthly on the policy documents is expensive, time-consuming, and requires storing and managing multiple model versions, which directly contradicts the requirement to avoid retraining. Option D is wrong because training a custom model from scratch each month is prohibitively expensive, requires vast amounts of data and compute resources, and is entirely unnecessary for a task that only needs to retrieve and synthesize existing information.

23
MCQeasy

Which machine learning paradigm involves training an agent to make decisions by interacting with an environment and receiving rewards or penalties based on its actions?

A.Unsupervised learning
B.Reinforcement learning
C.Supervised learning
D.Self-supervised learning
AnswerB

Reinforcement learning trains an agent through trial-and-error interaction with an environment, using reward and penalty signals to shape a policy. This directly matches the stem's requirement for decisions driven by rewards or penalties, unlike supervised or unsupervised paradigms that learn from static labelled or unlabelled datasets.

Why this answer

Reinforcement learning (RL) is the correct paradigm because it explicitly involves an agent learning a policy through trial-and-error interactions with an environment, receiving scalar reward signals (positive or negative) to maximize cumulative reward. This matches the question's description of making decisions based on rewards or penalties, which is the defining characteristic of RL, as opposed to learning from labeled data or discovering hidden patterns without feedback.

Exam trap

CompTIA AI often tests the distinction between reinforcement learning and supervised learning by phrasing the question to emphasize 'rewards or penalties' — candidates mistakenly think supervised learning uses penalties (like loss functions) and confuse it with RL's delayed reward signals from an environment.

How to eliminate wrong answers

Option A is wrong because unsupervised learning discovers hidden patterns or structures in unlabeled data without any reward or penalty signals from an environment. Option C is wrong because supervised learning maps inputs to outputs using labeled training data, where the model receives direct error feedback (e.g., loss function) rather than delayed rewards from environmental interactions. Option D is wrong because self-supervised learning generates its own supervisory signal from the input data itself (e.g., predicting masked tokens) and does not involve an agent acting in an environment to receive rewards or penalties.

24
Multi-Selecthard

A healthcare startup is building a diagnostic support system using a large language model. The system must provide accurate, evidence-based answers and avoid generating harmful or fabricated information. Which THREE techniques should be implemented to achieve this? (Choose 3)

Select 3 answers
A.Retrieval-Augmented Generation (RAG)
B.Disabling output filtering to speed up generation
C.Using chain-of-thought prompting for reasoning steps
D.Increasing the temperature parameter to encourage creativity
E.Fine-tuning on medical textbooks and guidelines
AnswersA, C, E

RAG retrieves relevant medical literature to ground responses.

Why this answer

Option A (Retrieval-Augmented Generation, RAG) is correct because grounding the LLM's responses in an external, authoritative medical knowledge base (e.g., PubMed, clinical guidelines) at inference time supplies verifiable evidence and sharply reduces hallucination compared to relying on parametric memory alone. Option C (chain-of-thought prompting) is correct because eliciting intermediate reasoning steps improves the model's accuracy on complex diagnostic questions and makes its conclusions auditable, which supports evidence-based clinical decision support. Option E (fine-tuning on medical textbooks and guidelines) is correct because domain-specific supervised fine-tuning adapts the model's weights to accurate, curated medical content and terminology, raising factual reliability for the healthcare domain.

Option B does not belong because disabling output filtering removes safety guardrails and increases the risk of harmful content, the opposite of the stated requirement. Option D does not belong because raising the temperature increases sampling randomness and creativity, which promotes fabricated or inconsistent answers rather than accurate, evidence-based ones.

Exam trap

AI0-001 often tests whether candidates confuse 'creativity' parameters like temperature with accuracy-enhancing techniques — higher temperature is a distractor that sounds like it improves output but actually worsens factual reliability.

25
MCQhard

A deep learning engineer is training a transformer model and notices that validation perplexity increases after a few epochs while training perplexity continues to decrease. Which of the following is the MOST likely cause?

A.The temperature parameter is set too high
B.The batch size is too small
C.The learning rate is too low
D.The model is overfitting the training data
AnswerD

Validation perplexity rising while training perplexity keeps falling is the classic divergence signature: the model memorises training-specific patterns rather than generalisable ones, so held-out performance degrades. That gap between the two curves is precisely what overfitting produces, making it the most likely cause here.

Why this answer

The described pattern—decreasing training perplexity alongside increasing validation perplexity—is the classic signature of overfitting. The model is memorizing the training data rather than learning generalizable patterns, causing its performance on unseen validation data to degrade after a certain point in training.

Exam trap

CompTIA AI often tests the distinction between optimization issues (like learning rate or batch size) and generalization issues (like overfitting), and the trap here is that candidates may confuse a rising validation loss with a learning rate that is too high, when in fact the divergence between training and validation metrics is the definitive clue for overfitting.

How to eliminate wrong answers

Option A is wrong because the temperature parameter controls the sharpness of the output probability distribution during inference (e.g., in softmax), not the training dynamics or the divergence between training and validation loss; a high temperature would make predictions more uniform, not cause overfitting. Option B is wrong because a batch size that is too small typically introduces high gradient variance and can slow convergence or cause instability, but it does not directly cause the specific pattern of training loss decreasing while validation loss increases—that is a hallmark of overfitting, not a batch-size issue. Option C is wrong because a learning rate that is too low would cause the model to converge very slowly or get stuck in a local minimum, but both training and validation perplexity would likely plateau or decrease together; it would not produce a divergence where training perplexity continues to drop while validation perplexity rises.

26
Multi-Selectmedium

A company is deploying a chatbot using a large language model. They want to mitigate the risk of prompt injection attacks. Which TWO measures should be implemented?

Select 2 answers
A.Implement input validation and sanitisation
B.Use a system prompt that strictly defines the chatbot's behavior
C.Fine-tune the model on safe conversational examples
D.Use a larger context window
E.Limit the maximum output token length
AnswersA, B

Input validation and sanitisation strip or escape injected instructions before they reach the model, blocking attempts to override system prompts. This constrains untrusted user input at the application boundary, mitigating prompt injection without altering the model itself.

Why this answer

Option A is correct because input validation and sanitisation directly strip or neutralise adversarial payloads (e.g., embedded instructions, delimiter-breaking characters, or encoded jailbreak strings) before they reach the LLM, reducing the attack surface for prompt injection. Option B is correct because a strict system prompt establishes immutable behavioural boundaries and instruction hierarchy, making it harder for user-supplied text to override the model's intended role or exfiltrate system instructions. Option C is not a reliable mitigation because fine-tuning on safe examples does not prevent novel injection payloads at inference time and can even be undone by adversarial prompting.

Option D is irrelevant, since a larger context window merely allows more tokens to be processed and does not filter or constrain malicious instructions. Option E only caps response size and does nothing to stop an injected prompt from altering the model's behaviour or leaking data.

Exam trap

CompTIA often tests the misconception that fine-tuning or output limits can prevent prompt injection, when in fact these measures do not address the root cause of untrusted input being processed as instructions.

27
Multi-Selecthard

A company is deploying an LLM-powered application that answers questions based on internal documents. They want to minimize prompt injection attacks where users trick the model into ignoring instructions. Which THREE measures should they implement? (Select THREE)

Select 3 answers
A.Use a system-level prompt that clearly defines allowed behavior and boundaries
B.Set temperature to 0.0 for all queries
C.Allow the model to execute any code from user prompts for flexibility
D.Implement a separate classifier to detect and block injection attempts
E.Sanitize user inputs to remove special tokens or injection patterns
AnswersA, D, E

A system-level prompt establishes persistent instructions that take precedence over user turns, defining permitted behaviour and boundaries. This constrains the model's response space so user input is less able to override the application's original directives.

Why this answer

Option A is correct because a system-level prompt that clearly defines allowed behavior and boundaries establishes the model's operating rules and helps it distinguish trusted instructions from untrusted user content, which is a core defense against prompt injection. Option D is correct because a separate classifier trained to detect and block injection attempts adds an independent detection layer that can catch malicious prompts before they reach the LLM, reducing the chance that a user overrides system instructions. Option E is correct because sanitizing user inputs to remove special tokens or known injection patterns prevents attackers from injecting delimiters, role markers, or instruction-like text that could confuse the model's instruction hierarchy.

Option B is not correct because setting temperature to 0.0 only makes output more deterministic and does not prevent prompt injection. Option C is not correct because allowing the model to execute arbitrary code from user prompts dramatically increases risk and is the opposite of a security control.

Exam trap

The AI0-001 exam often tests the misconception that reducing model temperature or randomness can mitigate security threats, when in fact temperature only affects output creativity, not instruction adherence or input safety.

28
MCQhard

A team is fine-tuning a BERT model for a document classification task. They notice the model achieves high F1 scores on the training set but low F1 on the validation set. Which regularization technique would be MOST effective?

A.L1 regularization
B.L2 regularization
C.Dropout
D.Reduce batch size
AnswerC

Dropout randomly deactivates neurons during each training pass, forcing the network to learn redundant, distributed representations rather than memorising training samples. This directly targets the overfitting gap between high training F1 and low validation F1 described in the stem, making it the most effective regulariser for fine-tuned BERT classification.

Why this answer

Dropout is the most effective regularization for transformer-based models like BERT because it randomly deactivates neurons during training, forcing the network to learn redundant representations and preventing co-adaptation. BERT already includes dropout layers in its architecture (attention dropout, hidden dropout), and increasing or tuning the dropout rate directly addresses overfitting on the training set.

Exam trap

AI0-001 often tests the misconception that L2 regularization is the go-to fix for overfitting in deep learning — candidates overlook that dropout is already embedded in transformer architectures and is the more targeted regularization technique for BERT fine-tuning.

How to eliminate wrong answers

Option A (L1 regularization) is wrong because L1 promotes sparsity in weights and is rarely used for deep transformer fine-tuning; it can destabilize training and does not target the co-adaptation that causes overfitting in BERT. Option B (L2 regularization) is wrong because while L2 (weight decay) helps, it is already applied via AdamW in most BERT fine-tuning recipes and is less effective than dropout at regularizing attention and feed-forward layers. Option D (Reduce batch size) is wrong because batch size affects optimization dynamics and gradient noise, not regularization directly; smaller batches can even increase variance and do not prevent overfitting.

29
MCQmedium

A data scientist is building a model to predict credit default using historical loan data. The dataset contains 100,000 records with 50 features, including income, debt-to-income ratio, and loan amount. The target variable is binary (default vs. no default). The goal is to maximize interpretability while maintaining high accuracy. Which algorithm is MOST appropriate?

A.Logistic regression
B.Random forest
C.Gradient boosting machine
D.Decision tree
AnswerA

Logistic regression produces coefficients whose sign and magnitude directly indicate each feature's contribution to default probability, giving the interpretability the goal demands. With 100,000 records and 50 features, it also achieves high accuracy on this binary target without sacrificing transparency.

Why this answer

Logistic regression is a linear model that is highly interpretable (coefficients indicate direction and magnitude of feature impact) and performs well on binary classification with many features, especially when the goal is to balance interpretability with accuracy. It avoids the overfitting risk of a single decision tree and the black-box nature of ensembles.

Exam trap

The trap is equating 'high accuracy' with complex models — candidates may pick random forest or GBM for accuracy, but the question explicitly prioritizes interpretability, making logistic regression the best fit.

How to eliminate wrong answers

Option B is wrong because random forest is an ensemble of trees that sacrifices interpretability for accuracy; it is harder to explain feature effects. Option C is wrong because gradient boosting machines are even more powerful but less interpretable and prone to overfitting without careful tuning. Option D is wrong because a single decision tree is interpretable but unstable and less accurate than logistic regression on this type of data, especially with 50 features.

30
MCQmedium

A data scientist is building a model to predict whether a transaction is fraudulent. The dataset has 99.9% legitimate transactions and 0.1% fraudulent ones. Which evaluation metric is MOST appropriate to assess model performance given this class imbalance?

A.BLEU score
B.Accuracy
C.F1-score
D.Perplexity
AnswerC

F1-score combines precision and recall into a single harmonic mean, so it penalises models that ignore the 0.1% fraudulent minority. Unlike accuracy, which reaches 99.9% by predicting "legitimate" always, F1-score reflects performance on the positive class, satisfying the stem's class-imbalance constraint.

Why this answer

With 99.9% legitimate transactions and only 0.1% fraudulent ones, accuracy would be misleadingly high (99.9%) even if the model never predicts fraud. The F1-score is the harmonic mean of precision and recall, making it robust to class imbalance by penalizing both false positives and false negatives. This makes it the most appropriate metric for evaluating fraud detection performance.

Exam trap

A common trap is that candidates default to accuracy as the universal metric, failing to recognize that in extreme class imbalance (e.g., 99.9% vs 0.1%), accuracy becomes meaningless and F1-score is the standard alternative.

How to eliminate wrong answers

Option A is wrong because BLEU score is a metric for evaluating machine translation quality by comparing n-gram overlap, not for binary classification or imbalanced datasets. Option B is wrong because accuracy is misleading in extreme class imbalance; a model that always predicts 'legitimate' would achieve 99.9% accuracy but fail to detect any fraud. Option D is wrong because perplexity is a metric used in language models to measure how well a probability distribution predicts a sample, not for evaluating classification performance on imbalanced data.

31
MCQmedium

A hospital's AI team is building a model that estimates a patient's 10-year risk of developing heart disease from 30 clinical and lifestyle variables. A cardiologist asks the team to explain why the model produced a high-risk score for a specific patient, because clinicians are legally required to justify their recommendations. The team needs a technique that assigns a numeric contribution to each input feature for that individual prediction. Which approach should the team use?

A.Confusion matrix analysis on a held-out validation set
B.SHAP (SHapley Additive exPlanations) values
C.k-fold cross-validation with stratified sampling
D.Global feature importance from a random forest's mean decrease in impurity
AnswerB

SHAP values come from cooperative game theory and distribute the prediction among the input features, giving each feature a signed numeric contribution for the single patient being explained. Because the sum of the SHAP values plus the base value reconstructs the model output, the cardiologist can see exactly how much variables such as blood pressure or cholesterol pushed the risk score up or down, which satisfies the need for a per-patient justification.

Why this answer

The clinician needs a per-instance explanation that quantifies how each input variable contributed to one patient's predicted risk. SHAP values satisfy this because they compute additive feature attributions grounded in Shapley values, and their sum reconstructs the model's output for that individual. Aggregate importance, confusion matrices, and cross-validation all describe model behavior across datasets or thresholds rather than explaining a single prediction.

Exam trap

The trap here is assuming that any feature importance output explains an individual prediction, when most built-in importance measures are global averages across the whole dataset.

32
MCQmedium

A data scientist is evaluating a binary classification model. The model achieves 95% accuracy on the test set, but the precision is 0.60 and recall is 0.55. The dataset has 90% negative class samples. Which metric should the team focus on to improve the model?

A.F1 score
B.Perplexity
C.BLEU score
D.Accuracy
AnswerA

With 90% negatives, accuracy is misleading because predicting the majority class alone scores 0.90. F1 is the harmonic mean of precision and recall, so optimising it directly penalises both false positives and false negatives, addressing the weak 0.60/0.55 balance.

Why this answer

With 90% negative class samples, accuracy is misleading because a trivial majority-class classifier would score 90%. The low precision (0.60) and recall (0.55) indicate the model struggles on the minority positive class. The F1 score, being the harmonic mean of precision and recall, is the right metric to optimise because it balances both concerns on imbalanced data.

Exam trap

AI0-001 often tests metric selection on imbalanced data, so candidates pick accuracy because it looks high, missing that it is the wrong optimisation target when the positive class is rare.

How to eliminate wrong answers

Option B is wrong because perplexity is a language-model metric measuring how well a probability distribution predicts a sample, not applicable to binary classification evaluation. Option C is wrong because BLEU score evaluates machine translation or text generation quality by n-gram overlap, unrelated to classification. Option D is wrong because accuracy is precisely the misleading metric here — with 90% negatives, high accuracy can coexist with poor minority-class performance, so focusing on it would not improve the model.

33
MCQhard

A developer is fine-tuning a large language model for a legal document summarization task. They notice that during training, the loss decreases rapidly in the first few epochs but then plateaus with high variance. Which hyperparameter adjustment is MOST likely to help stabilize training?

A.Add L1 regularization
B.Decrease the learning rate
C.Increase the batch size
D.Increase the number of epochs
AnswerB

A learning rate that is too high causes the optimiser to overshoot minima, producing the plateau with high variance seen after the initial rapid loss drop. Lowering it reduces update step size, letting the model settle into a smoother minimum and stabilising training.

Why this answer

A high-variance loss plateau after rapid initial convergence typically indicates that the learning rate is too large, causing the optimizer to overshoot the minima and oscillate. Decreasing the learning rate allows smaller, more stable weight updates, reducing variance and enabling smoother convergence.

Exam trap

CompTIA often tests the misconception that high variance in loss is always solved by increasing batch size or regularization, when in fact the immediate cause is often an overly aggressive learning rate that prevents convergence.

How to eliminate wrong answers

Option A is wrong because L1 regularization adds a penalty on the absolute magnitude of weights to induce sparsity, which does not directly address high variance in the loss curve during fine-tuning. Option C is wrong because increasing the batch size reduces gradient noise and can stabilize training, but the question describes high variance after a plateau, which is more directly tied to learning rate oscillations rather than batch size. Option D is wrong because increasing the number of epochs does not fix the underlying instability; it may even exacerbate overfitting or variance if the learning rate remains too high.

34
Multi-Selecteasy

A machine learning team is splitting a dataset for a binary classification problem. They want to ensure robust evaluation and avoid data leakage. Which TWO practices should they follow? (Choose 2)

Select 2 answers
A.Normalise the entire dataset before splitting
B.Split into training, validation, and test sets
C.Include validation data in the training set for more data
D.Shuffle the data before splitting
E.Use the same split for all experiments
AnswersB, D

Separate training, validation, and test sets keep hyperparameter tuning and final evaluation isolated from training data, preventing leakage. This satisfies the stem's robust-evaluation requirement, since the test set remains untouched until the model is frozen.

Why this answer

Option B is correct because splitting the dataset into separate training, validation, and test sets allows the model to be tuned on the validation set while the test set remains untouched for an unbiased final evaluation, which is essential for robust assessment of a binary classifier. Option D is correct because shuffling the data before splitting ensures that the training, validation, and test sets are representative of the overall distribution and prevents ordering bias or temporal artifacts from skewing the split, which helps avoid leakage when the data is not inherently ordered. Option A is not correct because normalising the entire dataset before splitting leaks information from the validation and test sets into the training process, since the scaling parameters would be computed using all data.

Option C is not correct because including validation data in the training set removes the independent validation signal needed for hyperparameter tuning and model selection, undermining robust evaluation. Option E is not correct because reusing the same split for all experiments can lead to overfitting to that particular split and does not provide the variability needed to assess generalization reliably.

35
Multi-Selectmedium

A team is training a deep learning model for image classification. They observe that training accuracy is high but validation accuracy is low, indicating overfitting. Which TWO techniques should they apply to reduce overfitting? (Select TWO)

Select 2 answers
A.Use L2 regularization
B.Increase learning rate
C.Add dropout layers
D.Increase the number of layers
E.Reduce training data size
AnswersA, C

L2 regularization adds a penalty on squared weight magnitudes to the loss function, constraining the model's capacity to memorise training samples. This reduces the gap between high training accuracy and low validation accuracy caused by overfitting.

Why this answer

Option A, using L2 regularization, is correct because adding a weight-decay penalty (e.g., lambda * sum of squared weights) to the loss function constrains the magnitude of the model's weights, discouraging the network from fitting noise in the training set and thereby reducing overfitting. Option C, adding dropout layers, is correct because dropout randomly deactivates a fraction of neurons (e.g., p=0.5) during each training step, preventing units from co-adapting and forcing the network to learn more robust, generalizable features. Option B, increasing the learning rate, is not appropriate because a larger learning rate typically causes unstable or divergent training rather than reducing overfitting.

Option D, increasing the number of layers, would raise model capacity and generally worsen overfitting. Option E, reducing training data size, would make overfitting more severe by giving the model even less data to generalize from.

Exam trap

CompTIA AI often tests the misconception that increasing model complexity (more layers or data reduction) helps generalization, when in fact these actions typically worsen overfitting.

36
Multi-Selecteasy

A company wants to build a system that can generate new product images for an online catalog. Which TWO generative AI approaches are most suitable?

Select 2 answers
A.Diffusion models
B.Variational autoencoders (VAEs)
C.Generative Adversarial Networks (GANs)
D.BERT-based model
E.GPT-style language model
AnswersA, C

Diffusion models generate high-fidelity images by iteratively denoising random noise, giving fine control over output quality and diversity. This suits catalog product imagery, where photorealistic, varied visuals are required, and they avoid the mode collapse and training instability that plague adversarial approaches.

Why this answer

Diffusion models (A) are correct because they generate high-fidelity, photorealistic images by iteratively denoising random noise, making them ideal for producing new product images for an online catalog. Generative Adversarial Networks (C) are also correct because a generator-discriminator pair can synthesize realistic product imagery and can be trained to match a catalog's visual style. Variational autoencoders (B) are not the best fit here because their outputs tend to be blurrier and less photorealistic than diffusion or GAN results.

BERT-based models (D) are encoder-only language models for understanding text, not image generation. GPT-style language models (E) generate text, not images, so they do not meet the requirement.

Exam trap

AI0-001 often tests the mapping between generative model families and output modalities — candidates confuse language models (BERT, GPT) with image generators, or assume any 'generative' model can produce images.

37
MCQmedium

A team is deploying a sentiment analysis model for social media posts. The model currently performs well on English text but poorly on code-switched text (e.g., Spanglish). Which approach is MOST effective for improving performance on code-switched data without starting from scratch?

A.Use a larger base model without additional training
B.Apply data augmentation by translating all code-switched posts to English
C.Train a new model from scratch on a mix of English and code-switched data
D.Fine-tune the existing model on a corpus of code-switched text
AnswerD

Fine-tuning adapts the existing model's weights to code-switched patterns, satisfying the requirement to avoid training from scratch. It teaches the model to handle intra-sentence language mixing, which the original English-only training data never exposed it to.

Why this answer

Fine-tuning the existing model on a corpus of code-switched text adapts the model to the new language pattern efficiently.

38
Multi-Selecteasy

A company wants to use machine learning to recommend products to customers based on their purchase history. Which TWO techniques are appropriate for this task? (Select TWO)

Select 2 answers
A.Collaborative filtering
B.Principal Component Analysis (PCA)
C.K-Nearest Neighbors (KNN)
D.Naive Bayes
E.Linear regression
AnswersA, C

Collaborative filtering recommends items by exploiting similarity patterns across users' purchase histories — users who bought similar items receive comparable suggestions. It requires no item content metadata, matching the stem's purchase-history-only input, and scales well for product recommendation.

Why this answer

Collaborative filtering (A) is correct because it is the canonical recommendation technique that leverages patterns across users' purchase histories (user-item interaction matrices) to predict products a customer is likely to buy, either via user-based or item-based similarity. K-Nearest Neighbors (C) is also correct because it can be applied to recommendation by finding the k most similar users or items based on historical purchase vectors and aggregating their preferences to generate recommendations. PCA (B) is a dimensionality-reduction technique, not a recommender, though it may be used as preprocessing.

Naive Bayes (D) is a probabilistic classifier for labeled categories such as spam detection, not suited to ranking product recommendations from purchase history. Linear regression (E) predicts a continuous numeric value and does not model user-item preference relationships needed for product recommendation.

39
MCQhard

A research team is training a deep learning model for image classification using a small dataset of 1,000 labeled images. They are concerned about overfitting. Which combination of regularisation techniques would be MOST effective?

A.Use early stopping without any other regularisation
B.Dropout with a rate of 0.5 and L2 regularisation
C.L1 regularisation and batch normalisation
D.Increase learning rate and use momentum
AnswerB

Dropout at 0.5 randomly deactivates half the units each pass, forcing redundant representations, while L2 regularisation penalises large weights. Combined, they constrain model capacity on the 1,000-image dataset, directly addressing the stem's overfitting concern more effectively than either alone.

Why this answer

Dropout with a rate of 0.5 randomly deactivates half of the neurons during each training step, forcing the network to learn redundant representations and reducing co-adaptation. L2 regularisation adds a penalty proportional to the square of weights to the loss function, discouraging large weights and smoothing the model. Together they combat overfitting from two complementary angles — architectural stochasticity and weight magnitude control — which is the most effective combination among the options.

Exam trap

AI0-001 often tests the misconception that any single regularisation technique is sufficient; the trap is choosing early stopping or batch normalisation as if they were equivalent to dropout plus weight decay, when the question asks for the most effective combination.

How to eliminate wrong answers

Option A is wrong because early stopping alone only halts training when validation performance degrades; it does not constrain model capacity or weight magnitude, so overfitting can still occur within the training run. Option C is wrong because L1 regularisation promotes sparsity (many weights become zero) but does not provide the stochastic regularization of dropout, and batch normalisation primarily stabilizes and accelerates training rather than directly preventing overfitting. Option D is wrong because increasing the learning rate and using momentum affects optimization dynamics, not regularization; a higher learning rate can even destabilize training and worsen generalization.

40
Multi-Selectmedium

A data scientist is preparing a dataset for a binary classification model. The dataset has 1000 samples, with 800 positives and 200 negatives. To evaluate the model properly, which THREE steps should they take? (Select THREE)

Select 3 answers
A.Remove the minority class samples to make the dataset balanced
B.Use a stratified train-test split to preserve class proportions
C.Apply SMOTE (Synthetic Minority Over-sampling Technique) to balance the training set
D.Report only accuracy as the evaluation metric
E.Use precision, recall, and F1-score for evaluation
AnswersB, C, E

A stratified split preserves the 80:20 positive-to-negative ratio in both training and test sets, satisfying the requirement for proper evaluation. Without stratification, random splitting could yield test sets with skewed class proportions, distorting precision and recall.

Why this answer

Option B is correct because a stratified train-test split preserves the 80/20 class ratio (800 positives, 200 negatives) in both the training and test sets, ensuring the evaluation reflects the true class distribution and avoids sampling bias. Option C is correct because SMOTE generates synthetic minority-class samples by interpolating between existing minority instances, balancing the training set so the classifier does not become biased toward the majority class. Option E is correct because with imbalanced data, accuracy is misleading; precision, recall, and F1-score reveal how well the model identifies the minority class and balances false positives against false negatives.

Option A is wrong because deleting minority samples discards valuable information and worsens the class imbalance problem. Option D is wrong because a model predicting all positives would achieve 80% accuracy while completely failing on the minority class, making accuracy an unreliable metric here.

Exam trap

CompTIA often tests the misconception that removing minority samples or relying solely on accuracy is acceptable for imbalanced datasets, when in fact these approaches degrade model performance and evaluation validity.

41
MCQeasy

A small logistics company wants to forecast next month's shipment volume using three years of historical monthly totals. The operations manager notes that volume has grown steadily and that December is always the busiest month. The data science consultant recommends a classical time series method that explicitly separates the long-term upward movement from the repeating yearly pattern. Which technique BEST fits this requirement?

A.k-means clustering on the monthly shipment totals
B.SARIMA (Seasonal AutoRegressive Integrated Moving Average)
C.A convolutional neural network trained on the monthly totals as a 1D signal
D.Logistic regression using the month number as the predictor
AnswerB

SARIMA extends ARIMA with seasonal autoregressive, differencing, and moving average terms, so it models trend and a repeating seasonal cycle simultaneously. With monthly data showing an upward trend and a strong December peak, the seasonal period of twelve captures the yearly pattern while the non-seasonal terms handle the growth. This makes SARIMA the direct match for the manager's stated need to separate trend from seasonality.

Why this answer

The scenario calls for decomposing a monthly series into a long-term trend and a twelve-month seasonal cycle, then projecting it forward. SARIMA is built precisely for that structure, using seasonal differencing and seasonal AR/MA terms alongside non-seasonal ones. Clustering has no forecasting capability, a CNN is impractical on thirty-six points and not interpretable as trend plus season, and logistic regression targets a categorical outcome rather than continuous volume.

Exam trap

The trap here is reaching for a flexible machine learning model by default, when classical time series methods like SARIMA are the correct tool for short series with clear trend and seasonality.

42
MCQhard

A team is training a deep learning model for image classification. The training loss decreases steadily but the validation loss plateaus after 20 epochs and then starts to increase. Which action is MOST likely to improve generalization?

A.Add more convolutional layers
B.Increase the learning rate
C.Implement early stopping
D.Reduce the batch size
AnswerC

Early stopping halts training at the epoch where validation loss begins rising, restoring the weights from the best validation checkpoint. This directly counters the stem's overfitting pattern, improving generalisation by preventing further memorisation of training data.

Why this answer

Early stopping halts training when validation loss stops improving, preventing overfitting. Increasing learning rate would worsen divergence; adding more layers increases capacity and overfitting; reducing batch size may help optimization but not directly address overfitting.

43
MCQmedium

A team is using a pre-trained BERT model for a sentiment analysis task on product reviews. They want to adapt it to their specific domain with limited labeled data. Which approach is MOST effective?

A.Use BERT as a feature extractor and train a logistic regression on top
B.Apply data augmentation to increase the dataset and then train from scratch
C.Train a new BERT model from scratch on the domain data
D.Fine-tune the pre-trained BERT model on the small labeled dataset
AnswerD

Fine-tuning updates BERT's pre-trained weights on the small labelled dataset, transferring general language representations to the sentiment domain. This suits limited labelled data far better than training from scratch, which would overfit the small sample.

Why this answer

Fine-tuning the pre-trained BERT model on the small labeled dataset is the most effective approach because BERT has already learned rich language representations from large-scale corpora. Fine-tuning updates all or some of the pre-trained weights on the target task, allowing the model to adapt to the domain with limited data. This transfer learning approach consistently outperforms feature extraction and training from scratch when labeled data is scarce.

Exam trap

AI0-001 often tests the misconception that feature extraction is equivalent to fine-tuning; candidates may choose the simpler feature-extraction approach, but the exam expects recognition that fine-tuning is more effective for domain adaptation with limited labeled data.

How to eliminate wrong answers

Option A is wrong because using BERT as a frozen feature extractor and training a logistic regression on top only leverages the pre-trained representations without adapting them to the domain, which typically yields lower performance than fine-tuning. Option B is wrong because training from scratch on augmented data still requires massive amounts of data and compute; augmentation cannot compensate for the lack of a large corpus, and training from scratch discards the benefits of pre-training. Option C is wrong because training a new BERT model from scratch on the domain data is infeasible with limited labeled data and would lead to severe overfitting and poor performance.

44
Multi-Selectmedium

A data scientist is building a recommendation system for an e-commerce platform. The dataset includes user purchase history, product descriptions, and user demographics. The goal is to recommend products that a user is likely to purchase. Which TWO techniques are most appropriate for this task? (Select TWO.)

Select 2 answers
A.Content-based filtering
B.Association rule mining
C.Linear regression
D.Anomaly detection
E.Collaborative filtering
AnswersA, E

Content-based filtering matches a user's past purchases against product description features, so it exploits the product descriptions in the dataset and recommends items similar to those the user already bought, without needing other users' data.

Why this answer

Content-based filtering (A) is correct because it recommends items by matching a user's past purchase history against product descriptions, which is exactly the item-attribute data available here. Collaborative filtering (E) is correct because it leverages patterns across many users' purchase histories to recommend products a similar user is likely to buy, directly addressing the recommendation goal. Together these are the two standard recommender-system techniques suited to user purchase history, product descriptions, and demographics.

Association rule mining (B) finds co-occurrence rules like market-basket pairs but does not personalize recommendations to a specific user. Linear regression (C) predicts a continuous numeric value and is not designed for ranking or recommending items. Anomaly detection (D) identifies outliers and is unrelated to generating product recommendations.

Exam trap

AI0-001 often tests the distinction between recommendation techniques and general ML algorithms, so candidates might incorrectly select linear regression or anomaly detection because they are familiar ML methods, but they do not address the personalization and ranking required for recommendation systems.

45
MCQeasy

Which of the following best describes the difference between narrow AI and general AI?

A.Narrow AI is designed for a specific task; general AI aims to perform any cognitive task a human can.
B.Narrow AI relies on supervised learning; general AI uses unsupervised learning exclusively.
C.Narrow AI requires large datasets; general AI can learn from few examples.
D.Narrow AI can perform any intellectual task; general AI is limited to specific tasks.
AnswerA

Narrow AI is scoped to one specific task, such as image classification or recommendation, and cannot transfer that capability elsewhere. General AI denotes a hypothetical system able to perform any cognitive task a human can, matching the stem's task-specificity versus broad-capability distinction.

Why this answer

Narrow AI specializes in one task, while general AI would possess human-like cognitive abilities across domains.

46
Multi-Selecthard

A financial services firm is designing an AI system to detect fraudulent transactions. The dataset is highly imbalanced, with fraud representing less than 0.1% of transactions. The team wants to build a model that reliably identifies fraud while minimizing false positives that inconvenience customers. Which TWO techniques are MOST appropriate to address the class imbalance and evaluation needs? (Choose two.)

Select 2 answers
A.Increase the number of layers in the neural network to improve capacity
B.Apply class weighting or resampling to give more importance to fraudulent transactions
C.Remove all non-fraud transactions to balance the dataset
D.Optimize only for accuracy to ensure overall correctness
E.Use precision-recall AUC instead of accuracy to evaluate model performance
AnswersB, E

Class weighting or resampling adjusts the training process so that the rare fraud class has more influence. This helps the model learn fraud patterns instead of defaulting to the majority class. Combined with appropriate evaluation, this technique directly addresses the imbalance and supports reliable fraud detection while allowing control over false positives through threshold tuning.

Why this answer

The two appropriate techniques are using precision-recall AUC for evaluation and applying class weighting or resampling. Precision-recall AUC highlights performance on the rare fraud class, while class weighting or resampling ensures the model learns from fraud examples. Together they address both the training imbalance and the need for meaningful evaluation, supporting reliable fraud detection with controlled false positives.

Exam trap

The trap here is prioritizing overall accuracy in an imbalanced fraud scenario, which rewards majority-class predictions and hides the model's failure to detect fraud.

47
MCQhard

A team is training a generative adversarial network (GAN) to generate realistic images of furniture. The generator loss decreases sharply while the discriminator loss increases. What is the MOST likely issue and recommended action?

A.Mode collapse has occurred; increase the generator's learning rate
B.The discriminator is overfitting; decrease its capacity
C.The learning rates are too high; reduce both
D.The generator is too strong; train the discriminator more frequently
AnswerD

When generator loss falls while discriminator loss rises, the discriminator can no longer distinguish real from fake, so the generator dominates. Training the discriminator more frequently restores adversarial balance, giving it enough updates to keep pace and prevent mode collapse.

Why this answer

When the generator loss decreases sharply while the discriminator loss increases, the generator is producing samples realistic enough to fool the discriminator consistently. This indicates the discriminator is not learning effectively, so the recommended action is to strengthen the discriminator by training it more frequently (or increasing its capacity). Training the discriminator more frequently gives it more opportunities to distinguish real from fake, restoring balance in the adversarial game.

Exam trap

The trap here is confusing generator dominance with mode collapse; candidates may pick mode collapse because it is a well-known GAN failure, but the described loss pattern points to discriminator weakness.

How to eliminate wrong answers

Option A is wrong because mode collapse is characterized by the generator producing limited variety of outputs, not by a sharp decrease in generator loss with increasing discriminator loss; increasing the generator's learning rate would likely worsen the imbalance. Option B is wrong because discriminator overfitting would typically show decreasing discriminator loss on training data but poor generalization, not an increasing discriminator loss. Option C is wrong because high learning rates would cause unstable, oscillating losses rather than a consistent divergence where the generator dominates.

48
MCQeasy

Which of the following is a key characteristic of Narrow AI (Weak AI)?

A.It can perform any intellectual task that a human can
B.It requires no training data
C.It is designed to excel at a single, specific task
D.It surpasses human intelligence in all domains
AnswerC

Narrow AI is engineered to perform one specific task, such as image classification or speech recognition, within a bounded domain. This matches the stem's characteristic, distinguishing it from general AI, which would transfer reasoning across unrelated tasks.

Why this answer

Narrow AI, also known as Weak AI, is designed and trained to perform a single, specific task with high proficiency, such as language translation, image recognition, or playing chess. It cannot generalize its intelligence to other domains, which distinguishes it from Artificial General Intelligence (AGI). Option C correctly captures this fundamental characteristic.

Exam trap

CompTIA often tests the distinction between Narrow AI and AGI, and the trap here is that candidates confuse 'narrow' with 'limited performance' rather than understanding it means 'restricted to a single task domain'.

How to eliminate wrong answers

Option A is wrong because the ability to perform any intellectual task that a human can describes Artificial General Intelligence (AGI), not Narrow AI, which is limited to a specific domain. Option B is wrong because Narrow AI systems require extensive training data to learn patterns and make accurate predictions; without training data, they cannot function. Option D is wrong because surpassing human intelligence in all domains is a trait of superintelligence, which is a theoretical concept beyond current Narrow AI capabilities.

49
MCQmedium

A team trains a neural network for image classification. During training, the loss decreases on the training set but increases on the validation set after a few epochs. What is the most likely cause?

A.Vanishing gradients
B.Incorrect learning rate scheduling
C.Overfitting
D.Underfitting
AnswerC

Overfitting occurs when the model memorises training data, including noise, so training loss keeps falling while validation loss rises after a few epochs. The diverging curves in the stem are the classic signature: the network has learned patterns that do not generalise to unseen data.

Why this answer

Overfitting occurs when the model learns the training data too well, including noise and irrelevant patterns, causing it to memorize rather than generalize. This is evidenced by the loss decreasing on the training set while increasing on the validation set after a few epochs, as the model's performance on unseen data degrades.

Exam trap

CompTIA AI often tests the distinction between overfitting and underfitting by presenting a scenario where training loss decreases but validation loss increases, which candidates may confuse with a learning rate issue or gradient problem.

How to eliminate wrong answers

Option A is wrong because vanishing gradients cause the network to stop learning entirely (loss plateaus on both sets), not a divergence between training and validation loss. Option B is wrong because incorrect learning rate scheduling typically causes erratic loss behavior (e.g., oscillations or failure to converge) on both sets, not a clear overfitting pattern. Option D is wrong because underfitting results in high loss on both training and validation sets, not a decreasing training loss with increasing validation loss.

50
Multi-Selecthard

A team is deploying a sentiment analysis model that must achieve high precision and high recall. They have a labeled dataset of 10,000 samples. They want to minimize overfitting. Which THREE actions are most appropriate? (Select THREE.)

Select 3 answers
A.Decrease the learning rate
B.Apply L2 regularization to the model weights
C.Use dropout layers in the neural network
D.Increase the training batch size
E.Augment the training data with synthetic examples
AnswersB, C, E

L2 regularization adds a penalty proportional to squared weights to the loss, shrinking coefficients and limiting the model's ability to fit training noise. On 10,000 sentiment samples this constrains variance, directly satisfying the stated goal of minimising overfitting while preserving precision and recall.

Why this answer

Option B is correct because applying L2 regularization adds a penalty proportional to the squared magnitude of the model weights to the loss function, which constrains weight growth and directly reduces overfitting, helping the model generalize and maintain both precision and recall on unseen data. Option C is correct because dropout layers randomly deactivate a fraction of neurons during each training iteration, preventing the network from relying on specific co-adapted units and acting as an effective regularizer that lowers overfitting. Option E is correct because augmenting the 10,000-sample training set with synthetic examples increases data diversity and effective sample size, which is a standard technique to improve generalization and reduce overfitting when labeled data is limited.

Option A is not appropriate because decreasing the learning rate only affects optimization step size and convergence stability, not the model's capacity to overfit, and can even slow convergence without regularizing. Option D is not appropriate because increasing the training batch size changes gradient estimation variance and training dynamics but does not by itself prevent overfitting; in fact, very large batches can sometimes harm generalization.

Exam trap

CompTIA often tests the misconception that decreasing the learning rate is a regularization technique, when in fact it only affects optimization speed and not model complexity or overfitting prevention.

51
MCQmedium

A company is deploying a large language model for customer support. They want to reduce the number of off-topic or nonsensical responses while maintaining creativity. Which parameter adjustment would BEST achieve this?

A.Decrease temperature to 0.2
B.Set top-p to 0.1
C.Increase top-k to 100
D.Increase temperature to 0.9
AnswerA

Lowering temperature to 0.2 sharpens the model's probability distribution, favouring high-likelihood tokens and suppressing erratic sampling. This directly curbs off-topic or nonsensical output while retaining some stochastic variation, satisfying the requirement to preserve creativity rather than collapsing to fully deterministic greedy decoding at temperature zero.

Why this answer

Lowering the temperature makes the model's probability distribution sharper, so it favors the highest-probability tokens and produces more deterministic, on-topic output. A value of 0.2 still allows some variation, preserving a degree of creativity while reducing off-topic or nonsensical responses.

Exam trap

AI0-001 often tests the direction of each sampling parameter — candidates confuse top-k/top-p (which widen or narrow the candidate pool) with temperature (which sharpens or flattens the distribution), and pick a top-p change when temperature is the intended lever.

How to eliminate wrong answers

Option B is wrong because setting top-p to 0.1 is an extremely aggressive nucleus sampling cutoff that can make responses overly rigid and may truncate valid continuations, and the question asks for the BEST adjustment to balance creativity with reduced off-topic output — temperature is the more direct and commonly recommended control. Option C is wrong because increasing top-k to 100 widens the candidate token pool, which increases randomness and off-topic risk rather than reducing it. Option D is wrong because increasing temperature to 0.9 flattens the distribution and increases randomness, producing more off-topic and nonsensical output, the opposite of the goal.

52
Multi-Selectmedium

A data scientist needs to select a regression model to predict house prices. The dataset contains many features, some of which are irrelevant. Which TWO algorithms are BEST suited for this scenario, and why? (Select TWO)

Select 2 answers
A.Ridge regression (L2 regularization)
B.Linear regression
C.Lasso regression (L1 regularization)
D.K-Nearest Neighbors
E.Random Forest
AnswersC, E

Lasso applies L1 regularisation, which drives irrelevant feature coefficients exactly to zero, performing automatic feature selection. This directly addresses the stem's many-features-with-some-irrelevant constraint, yielding a sparser, more interpretable model than ridge's L2 penalty, which shrinks but never eliminates coefficients.

Why this answer

Lasso regression (L1 regularization) is correct because its L1 penalty drives the coefficients of irrelevant features exactly to zero, performing automatic feature selection and yielding a sparse, interpretable model well suited to a dataset with many useless predictors. Random Forest is correct because its ensemble of decorrelated decision trees handles high-dimensional feature spaces robustly, captures non-linear relationships and interactions between features, and provides built-in feature-importance scores that tolerate irrelevant variables without overfitting as easily as a single model. Linear regression is not appropriate because it uses all features with no regularization, so irrelevant predictors inflate variance and degrade generalization.

Ridge regression shrinks coefficients toward zero but never eliminates them, so it does not perform feature selection. K-Nearest Neighbors is distance-based and suffers from the curse of dimensionality, making it a poor choice when many features are irrelevant.

Exam trap

The trap is treating Ridge and Lasso as interchangeable regularizers — candidates forget that only L1 (Lasso) produces sparse solutions that zero out irrelevant features, which is the key requirement in this scenario.

53
MCQeasy

A data scientist needs to predict whether a customer will churn (yes/no) based on historical data. Which type of machine learning problem is this?

A.Reinforcement learning
B.Regression
C.Binary classification
D.Clustering
AnswerC

Churn prediction produces one of two discrete outcomes, yes or no, so the target variable is binary. Binary classification is the problem type defined by exactly two mutually exclusive class labels, matching the stem's yes/no requirement.

Why this answer

This is a binary classification problem because the target variable has exactly two discrete outcomes: 'yes' (churn) or 'no' (no churn). Classification algorithms such as logistic regression, decision trees, or support vector machines are used to assign input features to one of these two predefined classes. The output is a categorical label, not a continuous value or a reward signal.

Exam trap

CompTIA often tests the distinction between classification and regression by presenting a binary outcome and expecting candidates to recognize it as classification, not regression, even though the term 'regression' appears in 'logistic regression' which is actually a classification algorithm.

How to eliminate wrong answers

Option A is wrong because reinforcement learning involves an agent learning to make sequences of decisions by interacting with an environment to maximize cumulative reward, not predicting a static binary outcome from historical data. Option B is wrong because regression predicts a continuous numeric value (e.g., revenue, temperature), not a discrete class label like churn yes/no. Option D is wrong because clustering is an unsupervised learning technique that groups data points based on similarity without using labeled target variables, whereas churn prediction requires labeled historical data to train a supervised model.

54
Multi-Selectmedium

An AI developer is selecting a model architecture for a real-time video surveillance system that must detect objects in each frame and also track movement patterns across frames. Which TWO architectures should the developer combine? (Choose 2)

Select 2 answers
A.Transformer encoder only
B.Generative adversarial network (GAN)
C.Variational autoencoder (VAE)
D.Recurrent neural network (RNN) or LSTM
E.Convolutional neural network (CNN)
AnswersD, E

RNNs and LSTMs maintain hidden state across timesteps, so they model temporal dependencies between frames — exactly the movement-tracking requirement. Combined with a CNN for per-frame object detection, they satisfy the stem's dual constraint of detecting objects and tracking motion patterns over time.

Why this answer

Option E, a convolutional neural network (CNN), is correct because CNNs apply learned spatial filters over pixel grids and are the standard architecture for per-frame object detection and feature extraction in video surveillance, efficiently capturing spatial hierarchies in each image. Option D, an RNN or LSTM, is correct because recurrent architectures model temporal dependencies across sequential frames, allowing the system to learn movement patterns and object trajectories over time; LSTMs in particular mitigate vanishing gradients for longer sequences. Together, a CNN front end for spatial detection plus an RNN/LSTM back end for temporal tracking forms a classic video-analysis pipeline.

Option A, a Transformer encoder only, is not the intended pairing here since it lacks the convolutional spatial inductive bias for frame-level detection and, used alone, does not provide the recurrent temporal modeling this scenario calls for. Option B, a GAN, is for generative adversarial training to synthesize or enhance data, not for detection and tracking. Option C, a VAE, is a generative model for learning latent representations and reconstruction, not a supervised detector-tracker component.

55
MCQmedium

A company is deploying a text generation model for customer service emails. They want to ensure the model's responses are factual and based on internal knowledge bases. Which technique is most effective?

A.Use Retrieval-Augmented Generation (RAG)
B.Fine-tune the model on historical customer service emails
C.Write a detailed system prompt
D.Set the temperature to 0
AnswerA

RAG retrieves relevant passages from internal knowledge bases at inference and conditions generation on them, so responses reflect actual company documentation rather than parametric guesses. This directly satisfies the factual, knowledge-base-grounded requirement for customer service emails.

Why this answer

Retrieval-Augmented Generation (RAG) is the most effective technique because it combines a pre-trained language model with a retrieval system that fetches relevant documents from an internal knowledge base at inference time. This ensures the model's responses are grounded in factual, up-to-date information from the knowledge base, reducing hallucinations.

Exam trap

AI0-001 often tests the difference between RAG and fine-tuning, and candidates may incorrectly choose fine-tuning for factual grounding when RAG is more appropriate.

How to eliminate wrong answers

Option B is wrong because fine-tuning on historical emails may improve style but does not guarantee factual accuracy or access to current knowledge; it can still hallucinate. Option C is wrong because a detailed system prompt can guide behavior but cannot provide the model with specific factual knowledge it lacks. Option D is wrong because setting temperature to 0 makes the model more deterministic but does not ensure factual correctness; it can still produce incorrect information.

56
Multi-Selecteasy

A company wants to classify images of products into categories. They have a large dataset of labeled images. Which TWO types of neural networks are most suitable for this task? (Select TWO.)

Select 2 answers
A.Generative Adversarial Network (GAN)
B.Convolutional Neural Network (CNN)
C.Recurrent Neural Network (RNN)
D.Transformer (e.g., Vision Transformer)
E.Multi-layer Perceptron (MLP)
AnswersB, D

Convolutional neural networks apply learned filters that exploit spatial locality and translation invariance in images, making them highly effective for classifying labelled product images. This satisfies the stem's requirement for a suitable architecture given a large labelled image dataset.

Why this answer

Option B, Convolutional Neural Network (CNN), is correct because CNNs use convolutional and pooling layers to exploit spatial locality and translation invariance in images, making them the classic architecture for supervised image classification with large labeled datasets. Option D, Transformer (e.g., Vision Transformer), is also correct because a Vision Transformer splits an image into patches, embeds them, and applies self-attention to model global relationships, achieving state-of-the-art results on image classification when sufficient labeled data (or pretraining) is available. Option A, Generative Adversarial Network (GAN), is not appropriate here because GANs are generative models that learn to synthesize data via a generator-discriminator game rather than directly performing supervised category classification.

Option C, Recurrent Neural Network (RNN), is unsuitable because RNNs are designed for sequential data such as text or time series and do not natively capture the 2D spatial structure of images. Option E, Multi-layer Perceptron (MLP), is not the best choice because fully connected layers ignore spatial hierarchy and typically underperform CNNs or Transformers on large-scale image classification.

Exam trap

The common mistake is assuming any general neural network can handle images, but RNNs and MLPs are not suited for spatial feature extraction despite being able to process image data.

57
MCQhard

A model trained on customer reviews achieves 98% accuracy on the test set. However, when deployed, it performs poorly on real-world data. The data scientist suspects distribution shift. Which action is MOST important to address this?

A.Reduce the learning rate during training
B.Implement a monitoring system to detect data drift and retrain with fresh data
C.Add more features to the model
D.Increase the number of cross-validation folds
AnswerB

Continuous monitoring detects when production input distributions diverge from training data, triggering retraining on fresh samples. This directly addresses distribution shift by closing the feedback loop between deployed predictions and current data, restoring performance that static evaluation on the original test set cannot capture.

Why this answer

Distribution shift (data drift) causes the model's training distribution to differ from the real-world distribution, degrading performance despite high test accuracy. Implementing a monitoring system to detect drift and retraining with fresh data directly addresses this by ensuring the model adapts to the current data distribution, which is the most critical action for maintaining performance in production.

Exam trap

CompTIA often tests the misconception that high test accuracy guarantees real-world performance, leading candidates to focus on training improvements (like tuning hyperparameters or adding features) rather than addressing the root cause of distribution shift through monitoring and retraining.

How to eliminate wrong answers

Option A is wrong because reducing the learning rate affects the optimization step size during training, which does not address distribution shift after deployment; it only changes how the model converges on the training data. Option C is wrong because adding more features may improve model capacity but does not fix the mismatch between training and real-world distributions; it could even exacerbate overfitting to the original distribution. Option D is wrong because increasing cross-validation folds improves the reliability of performance estimates on the training/validation data but does not detect or correct for distribution shift in the deployed environment.

58
MCQhard

An AI engineer is designing a system to detect unusual patterns in network traffic that may indicate a security breach. The system should learn from normal traffic patterns and flag deviations. Which machine learning approach is MOST appropriate?

A.Reinforcement learning with reward shaping
B.Supervised classification using logistic regression
C.Semi-supervised learning with a small labeled set
D.Unsupervised anomaly detection
AnswerD

Unsupervised anomaly detection learns the baseline distribution of normal traffic without labelled attack examples, then flags statistical deviations. This matches the scenario's constraint that breaches are unknown in advance, unlike supervised classification which needs labelled attack samples.

Why this answer

Unsupervised anomaly detection is the most appropriate approach because the system must learn 'normal' traffic patterns from unlabeled data and then flag deviations without requiring pre-labeled examples of attacks. This aligns with the core requirement of detecting unknown or novel security breaches, which supervised methods cannot handle due to the lack of labeled attack data.

Exam trap

The AI0-001 exam often tests the misconception that semi-supervised learning (Option C) is a middle ground for anomaly detection, but the trap is that it still requires labeled attack data, which is unavailable for unknown security breaches, making unsupervised methods the only viable choice.

How to eliminate wrong answers

Option A is wrong because reinforcement learning with reward shaping is designed for sequential decision-making problems (e.g., autonomous agents) and is not suited for static pattern detection in network traffic; it would require a reward function for 'normal' behavior, which is impractical for anomaly detection. Option B is wrong because supervised classification using logistic regression requires a fully labeled dataset of both normal and attack traffic, which is unavailable when the goal is to detect unknown or novel breaches. Option C is wrong because semi-supervised learning with a small labeled set still relies on labeled attack examples, which are scarce or nonexistent for novel security threats, and it does not purely model normal behavior like unsupervised methods do.

59
MCQeasy

A startup wants to identify unusual patterns in network traffic to detect potential security breaches. They have a large dataset of normal traffic but very few labeled attacks. Which machine learning approach is MOST suitable?

A.Supervised classification with logistic regression
B.Unsupervised anomaly detection
C.Reinforcement learning
D.Semi-supervised learning
AnswerB

With abundant normal traffic and almost no labelled attacks, supervised methods lack sufficient examples. Unsupervised anomaly detection learns the baseline distribution of normal traffic and flags statistical deviations, directly satisfying the constraint of scarce attack labels while surfacing unknown breach patterns.

Why this answer

Unsupervised anomaly detection is the most suitable approach because the startup has a large dataset of normal traffic but very few labeled attacks. This technique learns the baseline of normal behavior from unlabeled data and flags deviations as potential anomalies, which is ideal for detecting unknown or rare attack patterns without requiring labeled attack samples.

Exam trap

The AI0-001 exam often tests the misconception that semi-supervised learning is the best choice when labeled data is scarce, but the key distinction is that semi-supervised learning still requires a meaningful amount of labeled data for the target class, whereas unsupervised anomaly detection works with zero labeled attacks.

How to eliminate wrong answers

Option A is wrong because supervised classification with logistic regression requires a large, balanced set of labeled attack and normal traffic data to train effectively, which the startup lacks. Option C is wrong because reinforcement learning is designed for sequential decision-making problems (e.g., autonomous agents) and is not suited for static pattern detection in network traffic. Option D is wrong because semi-supervised learning still requires at least some labeled attack data to guide the model, and the startup has very few labeled attacks, making it less effective than pure unsupervised anomaly detection.

60
MCQeasy

An AI system that can perform any intellectual task that a human being can is referred to as:

A.Machine learning
B.Artificial General Intelligence (AGI)
C.Narrow AI
D.Deep learning
AnswerB

Artificial General Intelligence denotes a system matching human capability across any intellectual task, including reasoning, learning and transfer between domains. This contrasts with narrow AI, which is confined to a single specific task such as image classification.

Why this answer

Artificial General Intelligence (AGI) refers to a hypothetical AI system that possesses the ability to understand, learn, and apply intelligence across a wide range of tasks at a level comparable to a human being. It contrasts with narrow AI, which is specialized for specific tasks. The definition in the question matches AGI exactly.

Exam trap

AI0-001 often tests whether candidates confuse AGI with narrow AI or with specific techniques like machine learning and deep learning, which are subsets rather than categories of intelligence.

How to eliminate wrong answers

Option A is wrong because machine learning is a subset of AI focused on algorithms that learn from data, not a category of human-level general intelligence. Option C is wrong because Narrow AI (or Weak AI) is designed for a specific task, such as voice assistants or recommendation engines, and cannot perform any intellectual task a human can. Option D is wrong because deep learning is a technique using neural networks with many layers, not a classification of general intelligence.

61
MCQhard

A generative AI model is asked to 'Write a poem about AI' and returns a very short, generic response. The user wants longer, more creative outputs. Which parameter adjustment is MOST likely to help?

A.Decrease the top-p value
B.Increase the frequency penalty
C.Decrease the max tokens limit
D.Increase the temperature parameter
AnswerD

Higher temperature (e.g., 0.8-1.0) makes the model take more risks, leading to more creative and varied outputs.

Why this answer

Increasing the temperature parameter raises the randomness of token selection, encouraging the model to explore less probable word sequences and produce more varied, creative, and longer outputs. A low temperature (e.g., 0.1) makes the model deterministic and repetitive, often yielding short, generic responses. By increasing temperature (e.g., to 0.8 or 1.0), the model is more likely to generate diverse and expansive text, directly addressing the user's request for longer, more creative poems.

Exam trap

A common misconception in the CompTIA AI exam is that increasing max tokens (or decreasing it) is the primary way to control output length, when in fact temperature and top-p are the key parameters for influencing creativity and diversity, while max tokens simply sets a hard cutoff.

How to eliminate wrong answers

Option A is wrong because decreasing top-p (nucleus sampling) narrows the cumulative probability mass considered for token selection, making the output more focused and less diverse, which would further shorten and genericize the response. Option B is wrong because increasing the frequency penalty reduces the likelihood of repeating tokens or phrases, which can help with variety but does not directly encourage longer outputs; it may even shorten the response by penalizing common words. Option C is wrong because decreasing the max tokens limit explicitly caps the output length, which would make the response even shorter, opposite to the user's goal of longer outputs.

62
MCQmedium

A natural language processing team is building a system to classify support tickets into categories. They have a large corpus of unlabeled ticket text and a small set of manually labeled tickets. They want to leverage both to improve classification performance. Which approach is MOST suitable?

A.Apply unsupervised clustering to the unlabeled tickets and treat clusters as categories
B.Train a supervised classifier only on the small labeled set
C.Use semi-supervised learning that combines the unlabeled corpus with the labeled set
D.Perform principal component analysis on the ticket text and train a classifier on the components
AnswerC

Semi-supervised learning is designed for exactly this situation: a large amount of unlabeled data plus a small labeled set. Techniques such as self-training, co-training, or consistency regularization can use the unlabeled tickets to learn better representations and improve classification. This approach directly leverages both data sources, matching the team's objective and the data availability.

Why this answer

Semi-supervised learning is the best fit because it explicitly uses both the large unlabeled corpus and the small labeled set. Methods like self-training or consistency regularization can improve classification by learning from unlabeled tickets while still respecting the predefined categories. The other options either ignore one of the data sources or fail to produce a classifier aligned with the desired categories.

Exam trap

The trap here is assuming that any use of unlabeled data is clustering, when semi-supervised learning can incorporate unlabeled data while still training a supervised classifier for predefined categories.

63
MCQeasy

Which type of neural network is BEST suited for processing sequential data such as time series or natural language?

A.Generative Adversarial Network (GAN)
B.Multi-layer Perceptron (MLP)
C.Recurrent Neural Network (RNN)
D.Convolutional Neural Network (CNN)
AnswerC

RNNs process input sequentially, maintaining a hidden state that carries information from previous timesteps. This recurrent connection captures order and context in time series and natural language, where each element's meaning depends on those preceding it — the sequential-data constraint named in the stem.

Why this answer

Recurrent Neural Networks (RNNs) maintain a hidden state that carries information across time steps, making them inherently suited to sequential data like time series and natural language. Their recurrent connections allow them to model order and temporal dependencies, which feedforward architectures cannot do natively.

Exam trap

AI0-001 often tests the assumption that CNNs handle all data types, but candidates must recognize that sequential/temporal dependencies require the recurrent memory that only RNN-family architectures provide natively.

How to eliminate wrong answers

Option A is wrong because GANs are generative models consisting of a generator and discriminator used for creating synthetic data (images, audio), not for sequential sequence modeling. Option B is wrong because a Multi-layer Perceptron is a feedforward network with no memory of previous inputs, so it cannot capture temporal dependencies in sequences. Option D is wrong because CNNs are designed for spatial hierarchies (images) and, while 1D convolutions can process sequences, they lack the inherent recurrent memory that makes RNNs the canonical choice for sequential data.

64
MCQeasy

Which neural network architecture is specifically designed to process sequential data, such as time series or sentences, by maintaining a hidden state that captures information about previous inputs?

A.Transformer
B.Convolutional Neural Network (CNN)
C.Multi-layer Perceptron (MLP)
D.Recurrent Neural Network (RNN)
AnswerD

An RNN processes inputs sequentially, passing a hidden state forward at each timestep so earlier elements influence later outputs. This recurrent hidden state directly captures temporal dependencies in time series and word order in sentences.

Why this answer

Recurrent Neural Networks (RNNs) are specifically designed for sequential data because they maintain a hidden state that is updated at each time step, allowing information about previous inputs to persist and influence current and future outputs. This feedback loop makes them ideal for tasks like time series forecasting, natural language processing, and speech recognition, where order and context matter.

Exam trap

CompTIA often tests the misconception that Transformers are the default architecture for all sequence tasks, but the question specifically asks for a network that 'maintains a hidden state'—a defining feature of RNNs, not Transformers.

How to eliminate wrong answers

Option A (Transformer) is wrong because, while Transformers process sequences using self-attention mechanisms, they do not maintain a recurrent hidden state; they rely on positional encodings and parallel processing of the entire sequence. Option B (CNN) is wrong because CNNs are designed for spatial data (e.g., images) using convolutional filters and pooling layers, not for capturing temporal dependencies via a hidden state. Option C (MLP) is wrong because MLPs are feedforward networks with no memory or sequential processing capability; each input is processed independently without any hidden state carrying information across time steps.

65
MCQeasy

Which neural network architecture is specifically designed to handle sequential data and mitigate the vanishing gradient problem?

A.Convolutional Neural Network (CNN)
B.Transformer
C.Vanilla Recurrent Neural Network (RNN)
D.Long Short-Term Memory (LSTM) network
AnswerD

LSTM networks use gated cells (input, forget, output) that regulate information flow across time steps, preserving gradients over long sequences. This architecture directly mitigates the vanishing gradient problem while handling sequential data, matching the stem's requirement.

Why this answer

LSTM (Long Short-Term Memory) is a type of RNN designed with gating mechanisms to prevent vanishing gradients in long sequences. CNNs are for spatial data; vanilla RNNs suffer from vanishing gradients; transformers use attention but are not specifically designed to mitigate vanishing gradients (they use residual connections).

66
MCQhard

A computer vision engineer is building a model to detect defects on a manufacturing line. Defects are rare, occurring in only 0.5% of images. The engineer trains a convolutional neural network and achieves 99.5% accuracy, but the model never predicts a defect. The engineer wants to address the underlying issue. Which approach is MOST appropriate?

A.Switch from accuracy to mean squared error as the evaluation metric
B.Increase the learning rate to help the model escape the majority-class solution
C.Apply class weighting or resampling techniques to emphasize the minority defect class
D.Add more convolutional layers to increase model capacity
AnswerC

Class weighting or resampling directly counteracts the imbalance by increasing the cost of misclassifying defects or by presenting more defect examples during training. This forces the model to learn defect features instead of defaulting to the majority class. It is the standard, targeted remedy for a model that achieves high accuracy but fails on the rare class.

Why this answer

The model achieves high accuracy by always predicting the majority class, a classic symptom of severe class imbalance. Class weighting or resampling techniques adjust the training process so that minority-class errors carry more weight or appear more frequently, compelling the model to learn defect patterns. The other options do not target the imbalance and are unlikely to resolve the failure to detect rare defects.

Exam trap

The trap here is treating high accuracy as evidence of a good model when the metric is misleading under class imbalance, leading to fixes that ignore the need to rebalance the training signal.

Ready to test yourself?

Try a timed practice session using only Aio Ai Concepts Techniques questions.