Courseiva

CompTIA AI+ AI0-001 (AI0-001) — Questions 901–962

962 questions total · 13pages · All types, answers revealed

Page 12

Page 13 of 13

901
MCQeasy

An AI system that can perform any intellectual task that a human being can is referred to as:

A.Machine learning
B.Artificial General Intelligence (AGI)
C.Narrow AI
D.Deep learning
AnswerB

Artificial General Intelligence denotes a system matching human capability across any intellectual task, including reasoning, learning and transfer between domains. This contrasts with narrow AI, which is confined to a single specific task such as image classification.

Why this answer

Artificial General Intelligence (AGI) refers to a hypothetical AI system that possesses the ability to understand, learn, and apply intelligence across a wide range of tasks at a level comparable to a human being. It contrasts with narrow AI, which is specialized for specific tasks. The definition in the question matches AGI exactly.

Exam trap

AI0-001 often tests whether candidates confuse AGI with narrow AI or with specific techniques like machine learning and deep learning, which are subsets rather than categories of intelligence.

How to eliminate wrong answers

Option A is wrong because machine learning is a subset of AI focused on algorithms that learn from data, not a category of human-level general intelligence. Option C is wrong because Narrow AI (or Weak AI) is designed for a specific task, such as voice assistants or recommendation engines, and cannot perform any intellectual task a human can. Option D is wrong because deep learning is a technique using neural networks with many layers, not a classification of general intelligence.

902
MCQmedium

An AI system is being deployed to detect deepfakes in video content. To comply with transparency obligations, what should the company implement?

A.A system to automatically block all deepfake content
B.A visible or invisible watermark on AI-generated videos
C.A process to report deepfake content to law enforcement
D.Encryption of the video files to prevent tampering
AnswerB

Watermarking embeds a detectable marker in AI-generated video, letting viewers and automated systems identify synthetic content. This satisfies the transparency obligation by making the artificial origin of deepfake videos disclosable and verifiable, rather than relying on post-hoc detection alone.

Why this answer

Transparency obligations under AI governance frameworks (e.g., EU AI Act) require clear disclosure when content is AI-generated. A visible or invisible watermark directly informs viewers that the video is synthetic, fulfilling this requirement without over-blocking legitimate content. Option B is correct because it provides a verifiable, non-disruptive method of labeling AI-generated media.

Exam trap

CompTIA AI+ certification often tests the distinction between transparency (disclosure) and security (blocking, encryption, reporting), leading candidates to confuse a governance obligation with a technical control like blocking or encryption.

How to eliminate wrong answers

Option A is wrong because automatically blocking all deepfake content would violate freedom of expression and could suppress legitimate AI-generated art, satire, or educational material; transparency does not mandate censorship. Option C is wrong because reporting to law enforcement is a reactive, post-hoc measure that does not satisfy the proactive transparency obligation to label content at the point of consumption. Option D is wrong because encryption protects integrity and confidentiality but does not disclose the synthetic origin of the content to viewers, thus failing the transparency requirement.

903
Multi-Selecthard

Which THREE of the following are best practices for preventing overfitting in deep learning models?

Select 3 answers
A.L2 regularization
B.Increasing the number of layers
C.Dropout
D.Using a larger batch size
E.Data augmentation
AnswersA, C, E

L2 adds penalty on weights, keeping them small and reducing overfitting.

Why this answer

L2 regularization (also known as weight decay) adds a penalty term proportional to the square of the weight magnitudes to the loss function. This discourages the model from learning overly complex patterns by forcing weights to remain small, which reduces variance and helps prevent overfitting. It is a standard technique in deep learning frameworks like TensorFlow and PyTorch, where it is implemented via the `kernel_regularizer` or `weight_decay` parameter.

Exam trap

The AI0-001 exam often tests the misconception that increasing model complexity (e.g., more layers) or adjusting batch size are regularization techniques, when in fact they either worsen overfitting or serve different purposes like optimization speed.

904
MCQhard

A financial services firm runs an AI model that scores loan applications. Regulators require the firm to explain any adverse decision to an applicant. The model is a gradient-boosted tree with hundreds of features. Which implementation approach best satisfies the explainability requirement without replacing the model?

A.Replace the gradient-boosted tree with a logistic regression so coefficients are directly interpretable.
B.Publish the model's global feature importance ranking so applicants can see which factors matter most overall.
C.Generate per-applicant explanations using SHAP values that attribute the decision to the most influential features.
D.Log the model's raw input features and output score for each application so auditors can review the data.
AnswerC

SHAP values provide consistent, locally accurate feature attributions for individual predictions, which lets the firm explain why a specific applicant was declined. This satisfies the requirement to explain adverse decisions without discarding the high-performing tree model. The attributions can be presented in applicant-facing language and audited by regulators.

Why this answer

Adverse action explanations must be specific to the individual applicant, and SHAP values provide locally accurate attributions for each prediction. This preserves the gradient-boosted tree's performance while producing reasons that can be communicated and audited. Global importance describes the population, logistic regression replaces the model, and raw logs record inputs without explaining the decision.

Exam trap

The trap here is confusing global feature importance with per-applicant explanations, when adverse action notices require reasons tied to the individual decision.

905
MCQhard

An operations team runs a computer-vision model that flags manufacturing defects on an assembly line. Auditors require evidence that any single prediction can be reconstructed and explained months later. The team already logs model version, input image hash, and prediction score. Which additional logging practice best satisfies the audit requirement?

A.Log the prediction score together with the operator who reviewed the flagged unit.
B.Log the raw camera frames indefinitely and rely on the current model to regenerate explanations on demand.
C.Log the preprocessed feature vector, the model version identifier, and the explanation artifacts such as saliency or SHAP values for each inference.
D.Log only the aggregate daily defect rate and the model's overall precision.
AnswerC

Reconstructing and explaining one prediction requires the exact inputs the model consumed, the precise model version, and the attribution output that shows which features drove the score. Logging the feature vector plus explanation artifacts such as saliency maps or SHAP values gives auditors everything needed to replay and justify that specific decision months later.

Why this answer

Auditability of an individual prediction requires the exact inputs the model saw, the version of the model that scored them, and the explanation artifacts generated at inference time. Capturing the preprocessed feature vector, model version identifier, and attributions such as SHAP or saliency values lets auditors replay and justify any single decision. Aggregates, raw frames, or reviewer names cannot reconstruct the original reasoning.

Exam trap

The trap here is believing that storing raw inputs or aggregate metrics is enough for explainability, when the requirement is per-prediction attribution captured at inference time with the exact model version.

906
MCQeasy

Which AI accelerator is specifically designed by Google to accelerate the training and inference of large neural networks, especially in their cloud environment?

A.GPU
B.NPU
C.TPU
D.FPGA
AnswerC

Google-designed TPUs are application-specific integrated circuits built around systolic array matrix multiplication, which suits the massive parallel tensor operations in neural network training and inference. This directly satisfies the stem's requirement for a Google-designed accelerator operating within Google Cloud, unlike GPUs or general-purpose CPUs.

Why this answer

The Tensor Processing Unit (TPU) is Google's custom-designed ASIC specifically built to accelerate the training and inference of large neural networks. Unlike general-purpose hardware, TPUs are optimized for TensorFlow workloads and are a core component of Google Cloud's AI infrastructure, offering high throughput for matrix operations common in deep learning.

Exam trap

CompTIA often tests the distinction between custom-designed accelerators (like TPU) and general-purpose or reconfigurable hardware (like GPU, NPU, FPGA), expecting candidates to know that TPU is Google's proprietary solution for neural network acceleration in their cloud.

How to eliminate wrong answers

Option A is wrong because GPUs (Graphics Processing Units) are general-purpose parallel processors designed for graphics and compute, not specifically by Google for neural network acceleration in their cloud; they are widely used but not Google's custom accelerator. Option B is wrong because NPU (Neural Processing Unit) is a generic term for processors designed to accelerate neural networks, but it is not a specific Google-designed chip; Google's custom accelerator is the TPU. Option D is wrong because FPGAs (Field-Programmable Gate Arrays) are reconfigurable hardware that can be programmed for various tasks, but they are not specifically designed by Google for neural network training and inference in their cloud environment; Google uses TPUs for that purpose.

907
MCQmedium

A healthcare organization uses an AI model to predict patient readmission risk. To comply with patient privacy regulations, they apply differential privacy during training. What is the primary trade-off of using differential privacy?

A.Increased training time for reduced bias
B.Lower interpretability for higher fairness
C.Faster inference for lower memory usage
D.Reduced model accuracy for increased privacy
AnswerD

Differential privacy injects calibrated noise into training data or gradients, mathematically bounding any individual's influence on the model. This privacy guarantee inherently perturbs learned parameters, degrading predictive performance. The trade-off is therefore measurable accuracy loss, which the healthcare scenario accepts to satisfy patient privacy regulations.

Why this answer

Differential privacy works by adding calibrated noise to the training process or model outputs, which directly reduces the model's accuracy in exchange for a quantifiable privacy guarantee (e.g., ε-differential privacy). This trade-off is fundamental: stronger privacy (lower ε) requires more noise, which degrades predictive performance. The healthcare organization must balance the need to protect patient data against the clinical utility of accurate readmission predictions.

Exam trap

The AI0-001 exam often tests the misconception that differential privacy primarily reduces bias or improves fairness, when in fact its core trade-off is accuracy for privacy, and fairness can be negatively impacted by the added noise.

How to eliminate wrong answers

Option A is wrong because differential privacy does not primarily target bias reduction; it addresses privacy, and increased training time is a secondary implementation cost, not the primary trade-off. Option B is wrong because differential privacy does not inherently lower interpretability or increase fairness; it may even reduce fairness if noise disproportionately affects minority subgroups, and interpretability is a separate concern. Option C is wrong because differential privacy does not improve inference speed or reduce memory usage; it typically adds computational overhead during training and does not affect inference latency or memory footprint.

908
MCQhard

A hospital's AI triage assistant summarizes patient notes for clinicians. During post-deployment monitoring, the team notices the model's outputs drift in tone and length after the vendor silently updated the underlying foundation model. The application code did not change. Which action best restores reproducibility and protects against future silent model changes?

A.Increase the monitoring sample rate so the team detects drift faster on the next update.
B.Lower the max tokens parameter so outputs cannot drift in length after the vendor update.
C.Pin the model to a specific dated version or snapshot and re-run the evaluation suite before promoting any new version.
D.Add a post-processing step that rewrites every summary into a fixed template before display.
AnswerC

Pinning a dated model version or snapshot freezes behavior so results are reproducible, and re-running the evaluation suite before promotion catches regressions caused by vendor updates. This treats the foundation model as a versioned dependency, which is the standard way to control silent changes. It restores reproducibility without freezing the application or abandoning monitoring.

Why this answer

Silent foundation-model updates change behavior even when application code is untouched, so the model must be treated as a pinned, versioned dependency. Freezing a dated snapshot and requiring evaluation before promoting any new version restores reproducibility and blocks unevaluated changes from reaching clinicians. Token limits, output templates, and faster monitoring do not control which model version serves traffic.

Exam trap

The trap here is treating the foundation model as a stable service, when in fact vendor updates can change behavior without any application deployment.

909
MCQhard

A research team is training a deep neural network for image classification. The training loss decreases rapidly for the first few epochs but then plateaus, while validation loss starts to increase after epoch 10. Which action would best address this issue?

A.Reduce the batch size to introduce more noise during training.
B.Increase the learning rate to help the model escape the plateau.
C.Implement early stopping based on validation loss to prevent further overfitting.
D.Add more convolutional layers to increase model capacity.
AnswerC

Early stopping halts training when validation loss begins rising at epoch 10, while training loss keeps falling — the signature of overfitting. Restoring the best-performing weights prevents the network memorising training data, directly addressing the diverging loss curves described.

Why this answer

The training loss decreasing rapidly then plateauing while validation loss increases after epoch 10 is a classic sign of overfitting. Early stopping monitors validation loss and halts training when it begins to rise, preventing the model from memorizing noise in the training data. This directly addresses the overfitting issue without requiring architectural or hyperparameter changes that could destabilize training.

Exam trap

CompTIA often tests the misconception that plateauing training loss always requires adjusting learning rate or batch size, when in fact the simultaneous rise in validation loss is the definitive indicator of overfitting that early stopping is designed to solve.

How to eliminate wrong answers

Option A is wrong because reducing batch size increases gradient noise, which can actually worsen overfitting by preventing the model from converging to a stable minimum and may amplify validation loss increases. Option B is wrong because increasing the learning rate when validation loss is already rising risks overshooting the optimal weights, causing divergence or even higher validation loss. Option D is wrong because adding more convolutional layers increases model capacity, which exacerbates overfitting by giving the model more parameters to memorize training data rather than generalizing.

910
Multi-Selectmedium

An AI team is evaluating whether to use AI for a customer segmentation task. They have a dataset of customer demographics and purchase history. Which TWO conditions would make AI a better choice than a traditional rule-based approach? (Select two.)

Select 2 answers
A.The segmentation criteria are well-understood and can be expressed in simple if-then rules
B.The data contains complex, non-linear patterns that are not easily captured by rules
C.The business requires the model to adapt automatically as new customer data arrives
D.The segmentation must be fully explainable to regulators
E.The team has no access to labeled data
AnswersB, C

Why this answer

AI techniques like neural networks or gradient-boosted trees excel at capturing complex, non-linear interactions in high-dimensional data (e.g., purchase sequences combined with demographics) that rule-based systems cannot express without an explosion of brittle, hand-crafted conditions. This makes AI the better choice when the underlying patterns are not linearly separable or easily codified as if-then logic.

Exam trap

The AI0-001 exam often tests the misconception that AI is always superior to rule-based systems, but the trap here is that candidates overlook the specific constraints of explainability (Option D) and data requirements (Option E) that make rule-based approaches more appropriate in those contexts.

911
MCQhard

A financial institution uses a machine learning model to approve loan applications. The model was trained on historical data that inadvertently encoded a bias against applicants from certain zip codes, leading to discriminatory lending practices. A recent audit reveals that the model's decisions are unfair, and regulators require the bank to remediate the bias without significantly reducing overall approval accuracy. The data science team has access to the training data, the model, and a set of fairness metrics. They also have a small, unbiased validation set. Which course of action should the team take to satisfy regulatory requirements?

A.Remove the zip code feature from the model and retrain
B.Implement adversarial debiasing using the unbiased validation set to enforce fairness constraints
C.Increase the weight of samples from disadvantaged zip codes in the training data
D.Retrain the model using only the unbiased validation set
AnswerB

Adversarial debiasing directly optimizes for fairness and accuracy.

Why this answer

Adversarial debiasing directly addresses the bias encoded in the model by training a predictor and an adversary simultaneously. The adversary tries to predict the protected attribute (e.g., zip code) from the model's predictions, while the predictor is penalized for allowing such inference, enforcing fairness constraints. Using the unbiased validation set ensures the debiasing process is guided by ground truth labels that are free from historical bias, allowing the model to retain high accuracy while reducing discrimination.

Exam trap

CompTIA often tests the misconception that removing a sensitive feature (like zip code) is sufficient to eliminate bias, but the trap is that models can learn proxy features, so a more sophisticated debiasing technique like adversarial debiasing is required.

How to eliminate wrong answers

Option A is wrong because simply removing the zip code feature does not eliminate bias; the model can still learn proxy features (e.g., income, loan amount) that correlate with zip code, leading to continued discriminatory outcomes. Option C is wrong because increasing sample weights for disadvantaged zip codes may overcorrect and reduce overall accuracy, and it does not directly enforce a fairness constraint; it can also introduce new biases if the weighting is not carefully tuned. Option D is wrong because retraining on only the small unbiased validation set would likely lead to severe overfitting and poor generalization, as the dataset is too small to capture the full distribution of loan applications, significantly reducing approval accuracy.

912
MCQhard

A hospital's clinical decision support model was validated at 94 percent accuracy on a held-out set. After go-live, clinicians report that the model's suggestions are frequently irrelevant for elderly patients, even though overall accuracy in the monitoring dashboard has barely moved. Which monitoring practice would have surfaced this problem?

A.Measuring inference latency and throughput to confirm the model responds within clinical workflow time limits.
B.Tracking performance metrics disaggregated by patient demographic and clinical subgroups, with alerts when any subgroup falls below its own threshold.
C.Monitoring overall model accuracy and alerting when it drops more than two percentage points from the validation baseline.
D.Comparing the live distribution of input features against the training distribution to detect covariate shift.
AnswerB

A subgroup that is a small fraction of the population can degrade sharply while the aggregate metric stays flat, which matches the reported pattern. Segmenting metrics by age and other clinical attributes makes that hidden weakness visible. Subgroup thresholds then trigger action before clinicians lose trust in the system.

Why this answer

The clue is that aggregate accuracy held steady while a specific patient group received poor suggestions, which is the classic signature of a subgroup masked by population averaging. Disaggregated performance monitoring with per-segment thresholds is the practice that reveals it. Aggregate accuracy alerts, latency metrics, and input drift detection each observe a different dimension and cannot confirm that elderly patients specifically are being served badly.

Exam trap

The trap here is trusting a stable top-line accuracy number as proof of health, when a small or distinct subgroup can fail badly while the average barely moves.

913
MCQmedium

A company is developing a chatbot that helps users write code. They are concerned about the chatbot being used to generate malicious code. Which defense should they implement to reduce this risk?

A.Output filtering and guardrails to detect malicious code patterns
B.Input validation to block special characters
C.Data poisoning prevention during training
D.Red teaming the model before deployment
AnswerA

Output filtering inspects generated code before it reaches the user, blocking patterns matching malicious constructs such as reverse shells or credential harvesters. Guardrails therefore reduce the risk of the chatbot emitting harmful code, satisfying the stated concern.

Why this answer

Output filtering and guardrails inspect the model's generated response before it reaches the user, catching malicious code patterns such as reverse shells, credential stealers, or exploit payloads. This directly addresses the risk of the chatbot producing harmful code, regardless of what the user asked. Guardrails can combine pattern matching, classifiers, and policy rules to block or sanitize dangerous outputs.

Exam trap

The trap is choosing input validation because it sounds like the 'first line of defense' — but for code assistants, input filtering is both impractical (code needs special characters) and ineffective against a model that can generate harmful output from innocuous prompts.

How to eliminate wrong answers

Option B is wrong because input validation that blocks special characters would cripple a legitimate coding assistant — code inherently contains braces, semicolons, quotes, and symbols — and it does not stop the model from generating malicious code from benign-looking prompts. Option C is wrong because data poisoning prevention addresses tampering with training data, which is a training-time integrity concern, not the runtime risk of a user eliciting malicious code from an already-trained model. Option D is wrong because red teaming is a pre-deployment assessment activity that identifies weaknesses; it does not provide an ongoing runtime defense against malicious code generation, so it reduces risk only indirectly and not as an implemented control.

914
MCQmedium

A team is implementing a RAG system for legal document retrieval. The documents are long and cover multiple topics. Which chunking strategy is MOST appropriate to ensure each chunk contains coherent information?

A.Hierarchical chunking with overlapping windows
B.Semantic chunking based on topic boundaries
C.Fixed-size chunking with 512 tokens
D.Character-level chunking with no overlap
AnswerB

Semantic chunking splits where embedding similarity between adjacent sentences drops, so boundaries fall at genuine topic shifts rather than arbitrary token counts. Long, multi-topic legal documents therefore yield chunks that each stay within one coherent subject.

Why this answer

Semantic chunking based on topic boundaries is the most appropriate strategy because legal documents are long and cover multiple topics. By splitting at natural topic shifts (e.g., clauses, sections, or argument transitions), each chunk preserves coherent meaning, which is critical for accurate retrieval and generation in a RAG system. This approach avoids mixing unrelated content within a single chunk, which would degrade the quality of retrieved context.

Exam trap

CompTIA often tests the misconception that fixed-size token chunking is always optimal for simplicity, but in domain-specific RAG systems with long, multi-topic documents, semantic boundaries are essential to maintain chunk coherence and retrieval accuracy.

How to eliminate wrong answers

Option A is wrong because hierarchical chunking with overlapping windows adds complexity and redundancy without guaranteeing topic coherence; overlapping windows can introduce duplicate or fragmented information across chunks, which is inefficient for retrieval. Option C is wrong because fixed-size chunking with 512 tokens ignores semantic boundaries, often splitting a single legal argument or clause across two chunks, leading to incomplete or misleading context for the LLM. Option D is wrong because character-level chunking with no overlap destroys all semantic structure, producing arbitrary fragments that are useless for coherent retrieval and generation.

915
Multi-Selecthard

An organisation is deploying a fine-tuned LLM for internal use. They need to ensure the API endpoint is secure and cost-effective. Which TWO measures should they implement? (Choose 2)

Select 2 answers
A.Implement API key authentication
B.Enable content filtering
C.Disable logging to reduce storage costs
D.Apply rate limiting per user
E.Use gRPC instead of REST
AnswersA, D

API keys restrict access to authorised clients.

Why this answer

API key authentication (Option A) is a fundamental security measure that ensures only authorized clients can access the LLM endpoint. It provides a simple, lightweight mechanism to validate requests without the overhead of full OAuth, making it both secure and cost-effective for internal deployments.

Exam trap

The CompTIA AI+ exam tests the distinction between security measures (authentication, rate limiting) and non-security features (content filtering, protocol choice), leading candidates to mistakenly select content filtering or gRPC as security controls.

916
Multi-Selectmedium

A data scientist is using LIME to explain a black-box model. Which TWO characteristics of LIME are true?

Select 2 answers
A.It provides a measure of model confidence in its prediction
B.It requires access to the model's internal parameters
C.It provides a global ranking of feature importance across the entire dataset
D.It creates an interpretable surrogate model locally around a prediction
E.It can be used with any machine learning model
AnswersD, E

LIME fits an interpretable surrogate model, such as a sparse linear model, on samples generated locally around the instance being explained. This satisfies the stem's constraint because the approximation is faithful only near that prediction, not globally across the model.

Why this answer

Option D is correct because LIME (Local Interpretable Model-agnostic Explanations) works by perturbing the input around a specific prediction and fitting a simple, interpretable surrogate model (such as a linear model or decision tree) that approximates the black-box model's behavior in that local neighborhood. Option E is correct because LIME is model-agnostic: it only requires the ability to query the model's predictions (input-output access), so it can be applied to any machine learning model, including neural networks, SVMs, and gradient-boosted trees. Option A is not correct because LIME explains which features drove a prediction, not the model's confidence or probability calibration.

Option B is not correct because LIME treats the model as a black box and does not need internal parameters or gradients. Option C is not correct because LIME produces local explanations for individual predictions, not a global feature-importance ranking across the entire dataset.

Exam trap

AI0-001 often tests the distinction between local and global explanation methods, and candidates may mistakenly think LIME provides global feature importance or requires model internals.

917
MCQhard

A national security agency uses AI to analyze surveillance data for threat detection. The system is deployed in a high-stakes environment where false negatives could lead to missed threats, and false positives waste analyst time. Recently, a known hacker group attempted to evade detection by subtly modifying their communication patterns over time, a form of adversarial evasion. The agency wants to harden the system while maintaining performance. The system uses a deep neural network. Which mitigation strategy is most appropriate?

A.Switch to an unsupervised learning approach to detect anomalies
B.Simplify the model to a logistic regression to reduce the attack surface
C.Perform adversarial training using the hacker group's known evasion patterns
D.Add random noise to all input data to confuse evasion attempts
AnswerC

Adversarial training augments the training set with the group's known evasion patterns, so the deep neural network learns decision boundaries robust to those subtle modifications. This directly counters the observed evasion while retaining detection performance on legitimate traffic.

Why this answer

Adversarial training is the most appropriate mitigation because it directly incorporates known evasion patterns into the training process, making the deep neural network robust to the hacker group's subtle modifications. By retraining the model on adversarial examples, the decision boundary is hardened against these specific attacks without sacrificing overall detection performance. This approach is a standard defense in high-stakes security AI, balancing false positive and false negative rates while countering adversarial evasion.

Exam trap

CompTIA often tests the misconception that simplifying a model (e.g., to logistic regression) reduces attack surface, but in adversarial evasion, simpler models are actually more vulnerable because they lack the capacity to learn robust decision boundaries against crafted perturbations.

How to eliminate wrong answers

Option A is wrong because switching to unsupervised anomaly detection does not inherently defend against adversarial evasion; it may still be fooled by subtly modified patterns and often increases false positives due to lack of labeled threat data. Option B is wrong because simplifying to logistic regression reduces model capacity, making it less able to learn complex threat patterns and more susceptible to evasion, not less. Option D is wrong because adding random noise to input data degrades signal quality, increasing both false positives and false negatives, and does not target the specific evasion patterns used by the hacker group.

918
MCQhard

During a red-team exercise on an AI model, testers successfully extracted training data. Which vulnerability is this?

A.Membership inference
B.Model inversion
C.Adversarial example
D.Data poisoning
AnswerB

Model inversion reconstructs training-set samples by querying the model's outputs, directly satisfying the stem's constraint of extracted training data. Unlike membership inference, which only determines whether a record was used, inversion recovers actual feature values, making it the precise vulnerability demonstrated during this red-team exercise.

Why this answer

Model inversion attacks allow an adversary to reconstruct training data by exploiting the model's learned representations. In this scenario, the testers successfully extracted training data, which is the hallmark of a model inversion attack, not just inferring membership.

Exam trap

The AI0-001 exam often tests the distinction between 'extracting data' (model inversion) and 'inferring presence' (membership inference), so candidates mistakenly choose membership inference when the question explicitly states data was extracted.

How to eliminate wrong answers

Option A is wrong because membership inference only determines whether a specific data point was part of the training set, not extract the actual data. Option C is wrong because adversarial examples involve crafting inputs to cause misclassification, not extracting training data. Option D is wrong because data poisoning involves corrupting the training data to manipulate model behavior, not extracting existing training data.

919
MCQhard

An AI system misclassifies rare but critical events. The team considers using synthetic data. Which consideration is MOST important for ensuring the synthetic data improves performance on real rare events?

A.The synthetic data should include a wide variety of events, even if not realistic.
B.The synthetic data should be generated using an unsupervised generative model.
C.The synthetic data should accurately represent the distribution and features of real rare events.
D.The synthetic data should be as large as possible to cover all possibilities.
AnswerC

Synthetic samples only help if they mirror the true feature distribution and statistical properties of real rare events; otherwise the model learns artefacts that do not transfer, leaving the class-imbalance constraint unmet and real-world recall unchanged.

Why this answer

Synthetic data must faithfully replicate the distribution and feature space of real rare events to enable the model to learn meaningful decision boundaries. If the synthetic data does not capture the true underlying patterns—such as specific sensor readings or transaction anomalies—the model will fail to generalize to actual rare events, defeating the purpose of augmentation.

Exam trap

CompTIA often tests the misconception that 'more data is always better' or that 'any synthetic data helps,' when in reality the fidelity of the synthetic data to the real rare event distribution is the paramount factor for improving model performance on those events.

How to eliminate wrong answers

Option A is wrong because including a wide variety of unrealistic events introduces noise and spurious correlations, which can degrade the model's precision and recall on real rare events. Option B is wrong because the choice of generative model (unsupervised vs. supervised) is secondary; the critical factor is that the synthetic data accurately reflects the real rare event distribution, not the training paradigm. Option D is wrong because simply maximizing dataset size without ensuring fidelity to real rare events can lead to overfitting on synthetic artifacts and poor generalization to authentic edge cases.

920
MCQhard

A generative AI model is asked to 'Write a poem about AI' and returns a very short, generic response. The user wants longer, more creative outputs. Which parameter adjustment is MOST likely to help?

A.Decrease the top-p value
B.Increase the frequency penalty
C.Decrease the max tokens limit
D.Increase the temperature parameter
AnswerD

Higher temperature (e.g., 0.8-1.0) makes the model take more risks, leading to more creative and varied outputs.

Why this answer

Increasing the temperature parameter raises the randomness of token selection, encouraging the model to explore less probable word sequences and produce more varied, creative, and longer outputs. A low temperature (e.g., 0.1) makes the model deterministic and repetitive, often yielding short, generic responses. By increasing temperature (e.g., to 0.8 or 1.0), the model is more likely to generate diverse and expansive text, directly addressing the user's request for longer, more creative poems.

Exam trap

A common misconception in the CompTIA AI exam is that increasing max tokens (or decreasing it) is the primary way to control output length, when in fact temperature and top-p are the key parameters for influencing creativity and diversity, while max tokens simply sets a hard cutoff.

How to eliminate wrong answers

Option A is wrong because decreasing top-p (nucleus sampling) narrows the cumulative probability mass considered for token selection, making the output more focused and less diverse, which would further shorten and genericize the response. Option B is wrong because increasing the frequency penalty reduces the likelihood of repeating tokens or phrases, which can help with variety but does not directly encourage longer outputs; it may even shorten the response by penalizing common words. Option C is wrong because decreasing the max tokens limit explicitly caps the output length, which would make the response even shorter, opposite to the user's goal of longer outputs.

921
Multi-Selecteasy

A data scientist is tuning hyperparameters for a support vector machine (SVM) with an RBF kernel. Which two hyperparameters most significantly affect model performance? (Select TWO.)

Select 2 answers
A.gamma (kernel coefficient)
B.learning rate
C.epsilon (for epsilon-SVR)
D.degree (for polynomial kernel)
E.C (regularization parameter)
AnswersA, E

gamma determines the radius of influence of support vectors.

Why this answer

Gamma defines the influence of a single training example, with low values meaning a far reach and high values meaning a close reach. C controls the trade-off between achieving a low error on the training data and minimizing the margin, directly impacting overfitting. Together, they are the two most critical hyperparameters for an SVM with an RBF kernel.

Exam trap

Candidates often mistake kernel-specific hyperparameters (e.g., degree for polynomial, gamma for RBF) for general SVM parameters, selecting options like degree or epsilon without realizing they do not apply to the RBF kernel.

922
MCQeasy

During the data preparation phase of an AI project, a data scientist discovers that the target variable in a binary classification dataset is heavily imbalanced: 95% negative class and 5% positive class. Which technique should be applied to improve model performance on the minority class?

A.Apply oversampling of the minority class using techniques like SMOTE
B.Remove all samples from the majority class to balance the dataset
C.Normalize all numerical features to have zero mean and unit variance
D.Use a train-test split of 80-20 without any modification
AnswerA

SMOTE generates synthetic minority-class samples by interpolating between existing positive instances and their nearest neighbours, rebalancing the 95:5 distribution so the classifier no longer biases towards the majority class. This directly targets the minority-class performance constraint stated in the stem.

Why this answer

The dataset is heavily imbalanced (95% negative, 5% positive), which can cause models to be biased toward the majority class. Oversampling the minority class using techniques like SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic examples of the minority class, balancing the class distribution and improving the model's ability to learn the minority class patterns. This is a standard approach to address class imbalance.

Exam trap

AI0-001 often tests the misconception that accuracy is a good metric for imbalanced data, or that simple removal of majority class is acceptable; candidates might overlook the need for specialized resampling techniques.

How to eliminate wrong answers

Option B is wrong because removing all majority class samples would discard valuable information and likely lead to underfitting and poor generalization. Option C is wrong because normalization is a feature scaling technique that does not address class imbalance. Option D is wrong because a simple train-test split without addressing imbalance will still result in a model that performs poorly on the minority class.

923
MCQmedium

An operations team runs a demand-forecasting model on a cloud MLOps platform. The model retrains nightly, and after several weeks the live prediction distribution has drifted away from the distribution captured at training time. The team wants an automated signal that fires before prediction quality visibly degrades. Which practice should they implement?

A.Retrain the model on the full historical dataset each night and compare its validation accuracy to the previous release.
B.Increase the nightly retraining frequency to every hour so the model always reflects the newest data.
C.Configure drift detection on the model's input feature distributions and on the output prediction distribution, with alert thresholds tied to the training baseline.
D.Enable autoscaling of the inference endpoint so latency stays constant as request volume grows.
AnswerC

Monitoring the statistical distance between live feature and prediction distributions and the stored training baseline is exactly what produces an early warning before quality metrics like error rate degrade. Threshold alerts tied to that baseline turn the comparison into an actionable operational signal, which is what the team asked for.

Why this answer

The scenario calls for detecting a change in the relationship between live data and the training baseline before users notice worse predictions. Comparing current feature and prediction distributions against the stored training distributions, with thresholds that trigger alerts, gives that early warning. Retraining cadence, endpoint scaling, and periodic validation runs do not observe live distributional shift, so they cannot fire the needed signal.

Exam trap

The trap here is assuming that more frequent retraining or higher validation accuracy equals drift detection, when drift is about comparing live distributions to the training baseline.

924
MCQmedium

Refer to the exhibit. A data engineer runs a validation report on the customers table. The "income" column has 12 null values. Which imputation strategy is most appropriate for this column?

A.Remove rows with null income
B.Replace nulls with the median income per region
C.Replace nulls with 0
D.Replace nulls with the mean income of the entire dataset
AnswerB

Income is typically right-skewed and contains outliers, so the mean would be distorted; the median is robust. Computing it per region preserves local wage differences, satisfying the requirement to impute the 12 nulls without flattening genuine geographic variation in the column.

Why this answer

Imputing missing income values with the median per region preserves the central tendency of each regional subgroup, which is robust to outliers and maintains the distributional characteristics of the data. This strategy is particularly appropriate for income data, which often exhibits skewness and regional variation, ensuring that the imputed values are contextually relevant and do not distort downstream analytics or machine learning models.

Exam trap

CompTIA often tests the misconception that a global mean or median is always the safest imputation, when in fact ignoring subgroup structure (like region) can introduce significant bias and violate the assumption of missing-at-random conditioned on observed features.

How to eliminate wrong answers

Option A is wrong because removing rows with null income reduces the dataset size and can introduce bias if the missingness is not completely random, potentially degrading model performance and statistical power. Option C is wrong because replacing nulls with 0 is arbitrary and unrealistic for income data, introducing a strong downward bias that can severely skew summary statistics and mislead any analysis or model training. Option D is wrong because replacing nulls with the mean income of the entire dataset ignores regional heterogeneity and is sensitive to outliers, which can inflate variance and produce imputed values that are not representative of the local income distribution.

925
MCQhard

An MLOps engineer is deploying a scikit-learn random forest model to a Kubernetes cluster for a low-traffic internal API. The team wants to avoid maintaining a custom Flask wrapper and prefers a standard serving solution that supports REST and gRPC. Which serving component should they choose?

A.KServe with a scikit-learn model server runtime.
B.NVIDIA Triton Inference Server with a Python backend script that loads the pickle file.
C.TensorFlow Serving configured with a SavedModel export of the random forest.
D.A Kubernetes Deployment running a Flask app that loads the model and exposes only a REST endpoint.
AnswerA

KServe provides standardized model serving on Kubernetes with built-in support for REST and gRPC, autoscaling, and canary rollouts. Its scikit-learn runtime loads the pickled model directly without a custom Flask wrapper. This matches the requirement for a standard, low-maintenance serving solution for a random forest model on Kubernetes.

Why this answer

KServe offers a purpose-built scikit-learn runtime that serves pickled models over REST and gRPC on Kubernetes, with autoscaling and rollout features included. It eliminates the need for a bespoke Flask wrapper, which the team explicitly wants to avoid. The other options either require custom code, target GPU deep learning workloads, or demand unsupported format conversion.

Exam trap

The trap here is equating 'model serving on Kubernetes' with 'write a Flask container', ignoring that KServe provides ready-made runtimes for common frameworks like scikit-learn.

926
MCQhard

A team is building a natural language processing (NLP) model to analyze customer feedback. They have a large corpus of unlabeled text data and want to generate word embeddings that capture semantic meaning. Which approach should they use?

A.One-hot encoding
B.TF-IDF vectorization
C.Word2Vec
D.Bag-of-words model
AnswerC

Word2Vec learns dense vector representations from unlabelled text by predicting a word from its neighbours (skip-gram) or vice versa (CBOW), so semantic relationships emerge from co-occurrence statistics. This directly satisfies the stem's requirement for embeddings from a large unlabelled corpus, unlike supervised approaches needing labelled data.

Why this answer

Word2Vec is the correct approach because it learns dense, distributed word embeddings from large unlabeled corpora by training a shallow neural network to predict words in context (CBOW) or context from words (Skip-gram). This captures semantic relationships such as analogy and similarity, which is essential for analyzing customer feedback without labeled data.

Exam trap

CompTIA often tests the distinction between frequency-based vectorization (TF-IDF, bag-of-words) and prediction-based embedding methods (Word2Vec, GloVe), trapping candidates who think TF-IDF captures semantic meaning when it only captures term importance in a document.

How to eliminate wrong answers

Option A is wrong because one-hot encoding produces sparse, high-dimensional vectors with no semantic meaning—each word is represented as a binary vector with a single 1, and all vectors are orthogonal, so no similarity or relationship between words is captured. Option B is wrong because TF-IDF vectorization relies on term frequency and inverse document frequency to produce weighted sparse vectors, which reflect word importance in a document but do not capture semantic meaning or word relationships; it is a bag-of-words variant that ignores word order and context. Option D is wrong because the bag-of-words model creates sparse vectors based on word counts, losing all word order and context, and cannot generate embeddings that capture semantic similarity or analogy.

927
MCQmedium

A data engineer is building a pipeline to ingest streaming data from IoT sensors. Which data storage solution is best suited for real-time analytics on timestamped sensor readings?

A.Data warehouse
B.Relational database
C.Data lake
D.Time-series database
AnswerD

Time-series databases index on timestamp and exploit temporal locality, giving fast range scans and downsampling across sensor readings. This directly satisfies the real-time analytics constraint on timestamped IoT data, where relational stores would bottleneck on high-velocity writes and time-window aggregations.

Why this answer

Time-series databases (TSDBs) are optimized for high-ingest rates of timestamped data and provide efficient downsampling, retention policies, and time-based aggregation functions. For IoT sensor streaming, a TSDB like InfluxDB or TimescaleDB delivers sub-second query performance on time-range scans, which is essential for real-time analytics.

Exam trap

CompTIA often tests the misconception that 'any database can handle time-series data if you add a timestamp column,' ignoring the fundamental architectural differences in storage engines, indexing, and write optimization that make TSDBs the only viable choice for real-time streaming analytics.

How to eliminate wrong answers

Option A is wrong because data warehouses (e.g., Snowflake, Redshift) are designed for batch-oriented, structured querying of historical data and cannot sustain the high write throughput or low-latency time-range scans required for streaming sensor data. Option B is wrong because relational databases (e.g., PostgreSQL, MySQL) use row-based storage and B-tree indexes that degrade under continuous time-series inserts, leading to write contention and slow time-range queries. Option C is wrong because data lakes (e.g., S3, ADLS) store raw data in object storage with no indexing or time-ordering, making real-time analytics impossible due to high read latency and lack of native time-series functions.

928
MCQmedium

A retail company runs an AI-powered demand forecasting service in a Kubernetes cluster. The inference pods scale based on CPU utilization, but during flash sales the request queue grows rapidly and p99 latency spikes before new pods become ready. The operations team needs to reduce latency during these spikes without changing the model itself. Which action should the team take?

A.Enable a Kubernetes PodDisruptionBudget for the inference deployment and increase the replica count permanently.
B.Configure a Horizontal Pod Autoscaler (HPA) based on a custom external metric that reflects request queue depth, and lower the scale-up stabilization window.
C.Set the Kubernetes resource requests and limits for the inference pods to the maximum available node capacity.
D.Increase the model's batch size and enable request batching to improve throughput per pod.
AnswerB

Scaling on queue depth or concurrent requests reacts faster than CPU, which lags behind sudden bursts. Shortening the scale-up stabilization window lets the HPA add replicas sooner, reducing p99 latency during flash sales. This directly addresses the mismatch between the current CPU-based trigger and the actual bottleneck, the growing request queue.

Why this answer

The root cause is that CPU-based autoscaling reacts too slowly to a sudden queue buildup. Using a custom metric tied to queue depth or concurrent requests, combined with a shorter scale-up stabilization window, lets the HPA add capacity before latency degrades. The other options either add latency, waste resources, or do not improve reaction speed.

Exam trap

The trap here is assuming that any autoscaling configuration will fix latency, when the real issue is that the scaling signal does not reflect the actual bottleneck.

929
MCQhard

A platform team is preparing a feature store for a recommendation system. They need point-in-time correct feature retrieval so that training datasets do not leak future information, and they need the same features served online with low latency. Which architecture best satisfies both requirements?

A.Use a single relational database table that overwrites feature values in place with no timestamps.
B.Compute features on the fly in the training pipeline and again independently in the serving path using separate code.
C.Maintain an offline store with time-stamped feature values for training and a synchronized online store keyed by entity for low-latency serving.
D.Store features only in a data lake and have the online service query the lake on each request.
AnswerC

A dual-store feature platform keeps historical, time-stamped values in an offline store for point-in-time correct training joins, and materializes the latest values into a low-latency online store keyed by entity ID. This prevents label leakage during training while enabling fast online retrieval. It is the standard architecture for consistent offline and online features.

Why this answer

A feature store that pairs a time-stamped offline store with a synchronized online store supports point-in-time correct training joins and fast online retrieval from the same feature definitions. This dual-store pattern prevents label leakage and reduces training-serving skew. Single-store, lake-query, and duplicated-code approaches each fail at least one of correctness or latency.

Exam trap

The trap here is assuming that one storage system can serve both batch training and low-latency online inference, when the two access patterns require different stores synchronized through the same transformation logic.

930
MCQhard

An AI governance team is implementing the NIST AI Risk Management Framework. They have identified a high-risk AI system and are in the 'Measure' function. Which activity is most appropriate for this function?

A.Conduct bias and fairness impact assessments on the model
B.Implement technical controls to mitigate identified risks
C.Document the system's intended purpose and data sources
D.Establish an AI ethics board to oversee risk decisions
AnswerA

Bias and fairness impact assessments generate quantitative and qualitative evidence about model behaviour, which is exactly what the Measure function covers: analysing, benchmarking and monitoring identified risks. Mapping and framing belong to earlier functions, while this activity quantifies the high-risk system's performance.

Why this answer

In the NIST AI Risk Management Framework (AI RMF), the 'Measure' function focuses on assessing and analyzing risks associated with AI systems. For a high-risk AI system, conducting bias and fairness impact assessments is a core activity within this function, as it quantifies and evaluates potential harms related to fairness, accuracy, and transparency. This aligns with the framework's emphasis on quantitative and qualitative risk measurement before moving to risk treatment in the 'Manage' function.

Exam trap

The AI0-001 exam often tests the distinction between the NIST AI RMF functions (Map, Measure, Manage, Govern) by presenting risk mitigation actions (like implementing controls) as plausible activities for the 'Measure' function, when they actually belong to the 'Manage' function.

How to eliminate wrong answers

Option B is wrong because implementing technical controls to mitigate risks belongs to the 'Manage' function, which involves risk response and treatment, not the 'Measure' function that focuses on assessment and analysis. Option C is wrong because documenting the system's intended purpose and data sources is part of the 'Map' function, which establishes context and identifies risks, not the 'Measure' function that evaluates those risks. Option D is wrong because establishing an AI ethics board is a governance structure typically associated with the 'Govern' function, which sets policies and oversight, not the 'Measure' function's risk assessment activities.

931
MCQhard

A logistics company runs a vision model on edge devices in warehouses to detect damaged packages on conveyor belts. The model must classify each package within 40 milliseconds, and network connectivity to the cloud is unreliable. During a pilot, engineers notice that accuracy on the edge devices is several points lower than the accuracy measured during cloud-based evaluation on the same test images. Which cause is MOST likely?

A.Quantization applied to compress the model for the edge device reduced numerical precision in the weights and activations.
B.The edge devices have less RAM and slower CPUs than the cloud inference servers used during evaluation.
C.The test images were captured with a different camera model than the ones mounted on the conveyor belts.
D.The cloud evaluation used a batch size of 32 while the edge device processes one image at a time.
AnswerA

Converting a full-precision model to a smaller integer format for edge inference changes the numerical values the network computes, which commonly costs a few points of accuracy. Because the cloud evaluation used the original precision, the gap between the two environments points directly at the compression step. This is the classic source of an edge-versus-cloud accuracy delta on identical inputs.

Why this answer

When the same test images yield lower accuracy on the edge than in the cloud, the difference must come from something that changes the computation itself. Quantization to a lower-precision numeric format alters weights and activations, which is the standard cause of a small accuracy regression in compressed edge models. Camera mismatch would affect both environments equally, and hardware speed or batch size do not change the mathematical output for a given image.

Exam trap

The trap here is blaming hardware constraints or input differences for an accuracy gap, when the gap only appears after the model itself was transformed for edge deployment.

932
MCQmedium

A security analyst is evaluating adversarial threats to a deployed image classifier. Which attack involves making tiny, often imperceptible changes to input images to cause misclassification?

A.Model inversion
B.Membership inference
C.Adversarial examples
D.Data poisoning
AnswerC

Adversarial examples perturb input pixels by amounts imperceptible to humans, yet the cumulative gradient-aligned noise crosses the classifier's decision boundary, producing confident misclassification. This directly matches the stem's requirement for tiny input changes causing misclassification, unlike poisoning, evasion or model-inversion attacks, which alter training data or extract information instead.

Why this answer

Adversarial examples are inputs deliberately perturbed with small, often imperceptible changes to cause a machine learning model to misclassify. This matches the description of tiny changes to images leading to misclassification. The attack exploits the model's sensitivity to high-dimensional input spaces.

Exam trap

AI0-001 often tests the distinction between adversarial examples (inference-time input manipulation) and data poisoning (training-time data corruption), tempting candidates to confuse the two.

How to eliminate wrong answers

Option A is wrong because model inversion attempts to reconstruct training data from model outputs, not to cause misclassification via input perturbations. Option B is wrong because membership inference determines whether a specific record was in the training set, not to alter classification. Option D is wrong because data poisoning involves corrupting the training data to degrade model performance, not modifying inputs at inference time.

933
MCQmedium

A natural language processing team is building a system to classify support tickets into categories. They have a large corpus of unlabeled ticket text and a small set of manually labeled tickets. They want to leverage both to improve classification performance. Which approach is MOST suitable?

A.Apply unsupervised clustering to the unlabeled tickets and treat clusters as categories
B.Train a supervised classifier only on the small labeled set
C.Use semi-supervised learning that combines the unlabeled corpus with the labeled set
D.Perform principal component analysis on the ticket text and train a classifier on the components
AnswerC

Semi-supervised learning is designed for exactly this situation: a large amount of unlabeled data plus a small labeled set. Techniques such as self-training, co-training, or consistency regularization can use the unlabeled tickets to learn better representations and improve classification. This approach directly leverages both data sources, matching the team's objective and the data availability.

Why this answer

Semi-supervised learning is the best fit because it explicitly uses both the large unlabeled corpus and the small labeled set. Methods like self-training or consistency regularization can improve classification by learning from unlabeled tickets while still respecting the predefined categories. The other options either ignore one of the data sources or fail to produce a classifier aligned with the desired categories.

Exam trap

The trap here is assuming that any use of unlabeled data is clustering, when semi-supervised learning can incorporate unlabeled data while still training a supervised classifier for predefined categories.

934
MCQeasy

A company uses an AI system to generate marketing images. They are concerned about copyright ownership of the generated content. According to current US copyright law, who typically owns the copyright for AI-generated work?

A.No one; the work may be in the public domain
B.The user who provided the input prompts
C.The AI system itself, as the creator
D.The company that owns the AI model
AnswerA

US copyright law requires human authorship, so purely AI-generated images lack a protectable author. The work therefore falls into the public domain, meaning no party holds copyright, which is the position the scenario's ownership concern must account for.

Why this answer

Under current US copyright law and US Copyright Office guidance, copyright protection requires human authorship. Works generated entirely by AI without sufficient human creative input are not copyrightable and effectively fall into the public domain, meaning no one owns the copyright.

Exam trap

The trap is assuming the prompt author or the AI company automatically owns the output; the key legal principle is that copyright requires human authorship, so purely AI-generated works may have no owner.

How to eliminate wrong answers

Option B is wrong because providing input prompts alone has generally not been deemed sufficient human authorship by the Copyright Office; prompts are treated as instructions, not creative expression fixed in the output. Option C is wrong because an AI system is not a legal person and cannot hold copyright. Option D is wrong because owning the AI model does not confer copyright on outputs generated by users, absent a separate contractual assignment.

935
MCQeasy

Which type of neural network is BEST suited for processing sequential data such as time series or natural language?

A.Generative Adversarial Network (GAN)
B.Multi-layer Perceptron (MLP)
C.Recurrent Neural Network (RNN)
D.Convolutional Neural Network (CNN)
AnswerC

RNNs process input sequentially, maintaining a hidden state that carries information from previous timesteps. This recurrent connection captures order and context in time series and natural language, where each element's meaning depends on those preceding it — the sequential-data constraint named in the stem.

Why this answer

Recurrent Neural Networks (RNNs) maintain a hidden state that carries information across time steps, making them inherently suited to sequential data like time series and natural language. Their recurrent connections allow them to model order and temporal dependencies, which feedforward architectures cannot do natively.

Exam trap

AI0-001 often tests the assumption that CNNs handle all data types, but candidates must recognize that sequential/temporal dependencies require the recurrent memory that only RNN-family architectures provide natively.

How to eliminate wrong answers

Option A is wrong because GANs are generative models consisting of a generator and discriminator used for creating synthetic data (images, audio), not for sequential sequence modeling. Option B is wrong because a Multi-layer Perceptron is a feedforward network with no memory of previous inputs, so it cannot capture temporal dependencies in sequences. Option D is wrong because CNNs are designed for spatial hierarchies (images) and, while 1D convolutions can process sequences, they lack the inherent recurrent memory that makes RNNs the canonical choice for sequential data.

936
MCQeasy

Which neural network architecture is specifically designed to process sequential data, such as time series or sentences, by maintaining a hidden state that captures information about previous inputs?

A.Transformer
B.Convolutional Neural Network (CNN)
C.Multi-layer Perceptron (MLP)
D.Recurrent Neural Network (RNN)
AnswerD

An RNN processes inputs sequentially, passing a hidden state forward at each timestep so earlier elements influence later outputs. This recurrent hidden state directly captures temporal dependencies in time series and word order in sentences.

Why this answer

Recurrent Neural Networks (RNNs) are specifically designed for sequential data because they maintain a hidden state that is updated at each time step, allowing information about previous inputs to persist and influence current and future outputs. This feedback loop makes them ideal for tasks like time series forecasting, natural language processing, and speech recognition, where order and context matter.

Exam trap

CompTIA often tests the misconception that Transformers are the default architecture for all sequence tasks, but the question specifically asks for a network that 'maintains a hidden state'—a defining feature of RNNs, not Transformers.

How to eliminate wrong answers

Option A (Transformer) is wrong because, while Transformers process sequences using self-attention mechanisms, they do not maintain a recurrent hidden state; they rely on positional encodings and parallel processing of the entire sequence. Option B (CNN) is wrong because CNNs are designed for spatial data (e.g., images) using convolutional filters and pooling layers, not for capturing temporal dependencies via a hidden state. Option C (MLP) is wrong because MLPs are feedforward networks with no memory or sequential processing capability; each input is processed independently without any hidden state carrying information across time steps.

937
MCQmedium

A team is building a recommendation system for an e-commerce platform. They want to use collaborative filtering but have a cold-start problem for new users. Which hybrid approach BEST addresses cold start while leveraging collaborative signals?

A.Apply user clustering based on demographic data and then use collaborative filtering within clusters
B.Use only content-based filtering for all users
C.Use matrix factorization with implicit feedback only
D.Implement a hybrid model that combines content-based features with collaborative filtering via a weighted ensemble
AnswerD

A weighted ensemble blends content-based features, which describe new users from available attributes, with collaborative filtering signals from existing users. This supplies recommendations despite missing interaction history, directly resolving the cold-start constraint while preserving collaborative signal use.

Why this answer

A weighted ensemble hybrid combines content-based features (which can profile a brand-new user from their stated preferences, demographics, or first-item interactions) with collaborative filtering signals (which capture behavioral patterns from similar users). This directly mitigates cold start because the content-based component can generate recommendations before enough interaction data exists for collaborative filtering, while the collaborative component takes over as behavioral data accumulates. Pure collaborative filtering fails for new users because there is no interaction history to compute similarities.

Exam trap

AI0-001 often tests the misconception that clustering or demographic segmentation alone solves cold start, when in fact any collaborative method still requires interaction data that new users lack.

How to eliminate wrong answers

Option A is wrong because clustering by demographics and then applying collaborative filtering still requires interaction data within each cluster — a new user with no ratings cannot be matched to neighbors, so cold start persists. Option B is wrong because using only content-based filtering discards collaborative signals entirely, sacrificing the personalization quality that comes from user-item interaction patterns and failing to 'leverage collaborative signals' as the question requires. Option C is wrong because matrix factorization with implicit feedback still needs user-item interaction events to factorize; a new user with zero interactions produces an undefined or random latent vector, so cold start is not solved.

938
MCQmedium

A bank uses an AI system for credit scoring. To meet fairness requirements, they want to ensure the model predicts similar outcomes for individuals who are similar with respect to the target variable, regardless of protected attributes. Which fairness metric addresses this?

A.Demographic parity
B.Calibration
C.Individual fairness
D.Equalised odds
AnswerC

Individual fairness requires that similar individuals receive similar predictions, measuring outcome consistency between comparable cases irrespective of protected attributes. This matches the bank's requirement, unlike group fairness metrics such as demographic parity or equalised odds.

Why this answer

Individual fairness requires that similar individuals receive similar predictions — that is, the model treats people who are alike with respect to the target variable in the same way, regardless of protected attributes. This matches the bank's requirement to predict similar outcomes for similar individuals irrespective of protected characteristics.

Exam trap

AI0-001 often tests the distinction between group-level fairness metrics (demographic parity, equalised odds, calibration) and individual-level fairness, so candidates may pick a group metric when the question emphasises 'similar individuals'.

How to eliminate wrong answers

Option A is wrong because demographic parity requires equal positive prediction rates across groups, which is a group-level statistical criterion, not a similarity-based individual one. Option B is wrong because calibration concerns whether predicted probabilities match observed frequencies within groups, not whether similar individuals get similar outcomes. Option D is wrong because equalised odds requires equal true positive and false positive rates across groups, again a group-level criterion rather than an individual similarity guarantee.

939
MCQhard

A media company uses a reinforcement learning agent to schedule promotional banners on its homepage. The agent receives a reward when users click a banner, and it has learned to show the same sensational headline repeatedly because it historically generated high clicks. The editorial team is concerned that this harms long-term user trust. Which modification best aligns the agent's objective with long-term user satisfaction?

A.Increase the discount factor so future rewards are weighted more heavily.
B.Reduce the size of the experience replay buffer so the agent forgets older click patterns.
C.Switch from an epsilon-greedy policy to a softmax exploration policy.
D.Add a reward penalty for showing the same banner repeatedly and include a signal for user retention or satisfaction.
AnswerD

Redefining the reward to penalize repetition and include retention or satisfaction directly changes what the agent optimizes. This aligns the objective with long-term trust rather than short-term clicks. By incorporating both a repetition penalty and a satisfaction signal, the agent learns to balance engagement with sustainable user experience, which is exactly what the editorial team needs.

Why this answer

The agent's behavior stems from a reward function that values clicks without regard for long-term consequences. Adding a repetition penalty and a retention or satisfaction signal reshapes the objective so the agent learns to avoid overexposing sensational content and to prioritize sustainable user engagement.

Exam trap

The trap here is assuming that changing exploration or discount parameters will fix a misaligned objective, when the core issue is that the reward function itself rewards repetitive clickbait.

940
MCQhard

A data science team is deploying a real-time fraud detection model on edge devices in retail stores. The model must infer under 10ms and fit within 50MB memory. Which combination of techniques should the team apply?

A.Model parallelism and distributed inference
B.Increase batch size and use FP16 precision
C.Train a larger model and use distillation to transfer knowledge
D.Model quantization to INT8 and pruning of low-weight connections
AnswerD

Quantisation to INT8 shrinks weights from 32-bit to 8-bit, cutting memory roughly fourfold, while pruning removes low-weight connections to reduce computation. Together they meet the 50MB footprint and sub-10ms inference latency constraints on constrained edge hardware.

Why this answer

INT8 quantization reduces each weight from 32-bit float to 8-bit integer, cutting model size ~4x and enabling faster integer arithmetic that meets the sub-10ms latency target. Pruning removes low-magnitude weight connections, further shrinking the 50MB footprint and reducing compute. Together they are the standard edge-optimization pairing for latency- and memory-constrained inference.

Exam trap

AI0-001 often tests the misconception that more hardware parallelism or larger batches solve latency problems, when edge constraints actually demand model compression techniques like quantization and pruning.

How to eliminate wrong answers

Option A is wrong because model parallelism splits a model across multiple devices and distributed inference adds network hops — both increase latency and assume multiple nodes, which is impractical on a single constrained edge device. Option B is wrong because increasing batch size raises per-inference latency and memory, and FP16 alone only halves precision without the 4x reduction INT8 provides, so it won't reliably fit 50MB. Option C is wrong because training a larger model then distilling still yields a model that must be quantized/pruned to meet the constraints; distillation alone doesn't guarantee the 50MB/10ms targets.

941
MCQhard

A bank operates a credit-scoring model in production. Auditors require the team to reproduce the exact score a specific applicant received six months ago, including the model version, the feature values, and the code path used. Which capability must the team have in place to satisfy this requirement?

A.Full lineage logging that captures the model artifact version, the raw input record, the engineered feature values, and the inference request metadata for every prediction.
B.A model registry that stores every trained model version with its performance metrics and promotion status.
C.Periodic shadow deployment of the current model against the previous version to compare scoring behavior.
D.A dashboard that tracks aggregate approval rates and score distributions over time to demonstrate stable model behavior.
AnswerA

Reproducing an individual historical decision requires the exact artifact, the exact inputs after feature engineering, and the context of the request. Lineage logging that binds these together at inference time is the only listed capability that lets an auditor replay the decision deterministically. Without stored feature values, even the right model version cannot regenerate the same score.

Why this answer

Auditability of an individual prediction demands that the model version, the post-engineering feature values, and the request context all be captured at the moment of inference. Only comprehensive lineage logging binds those elements to a specific decision so it can be replayed later. Registries, aggregate dashboards, and shadow deployments each cover a different concern and none retains the record-level detail an auditor needs to reconstruct one score.

Exam trap

The trap here is treating a model registry as sufficient for auditability, when registry entries lack the per-request inputs and engineered features needed to replay a single decision.

942
Multi-Selecteasy

Which TWO of the following are best practices for securing an AI model against adversarial attacks?

Select 2 answers
A.Model pruning to reduce the number of parameters.
B.Adversarial training with perturbed examples.
C.Input sanitization and validation.
D.Increasing model complexity to capture more patterns.
E.Hyperparameter optimization using grid search.
AnswersB, C

Adversarial training with perturbed examples hardens the model by exposing it to manipulated inputs during fitting, so it learns decision boundaries robust to small, deliberate perturbations. This directly satisfies the stem's requirement for a best practise against adversarial attacks, reducing misclassification of crafted inputs at inference time.

Why this answer

Option B is correct because adversarial training explicitly augments the training set with perturbed inputs (e.g., FGSM or PGD examples) labeled with their true classes, which hardens the model's decision boundaries against small, intentionally crafted perturbations. Option C is correct because input sanitization and validation filter or normalize anomalous inputs (e.g., clipping pixel ranges, rejecting out-of-distribution values, or detecting unusual feature patterns) before inference, reducing the attack surface for adversarial examples. Option A is not a security best practice for adversarial robustness; pruning reduces parameters for efficiency and can even increase vulnerability by removing redundant features that aid generalization.

Option D is wrong because increasing model complexity typically enlarges the attack surface and can worsen overfitting to non-robust features, making adversarial examples easier to find. Option E is irrelevant to adversarial security, as grid search only tunes hyperparameters for performance metrics like accuracy, not robustness against crafted perturbations.

Exam trap

CompTIA often tests the misconception that increasing model complexity or pruning improves security, when in fact these techniques address performance or efficiency, not adversarial robustness.

943
MCQmedium

An AI model is trained to predict loan default. The training data contains 95% non-default and 5% default. Which metric is most appropriate to evaluate model performance given the imbalanced dataset?

A.Mean squared error
B.F1-score
C.Accuracy
D.R-squared
AnswerB

With only 5% defaults, accuracy misleads by favouring the majority class. F1-score is the harmonic mean of precision and recall, so it penalises both false positives and false negatives, giving a balanced view of performance on the minority default class.

Why this answer

The F1-score is the harmonic mean of precision and recall, making it robust to class imbalance. In this dataset with 95% non-default and 5% default, accuracy would be misleadingly high (95%) even if the model never predicts default, while F1-score penalizes poor recall of the minority class.

Exam trap

CompTIA often tests the misconception that accuracy is always the best metric, leading candidates to overlook its failure in imbalanced scenarios where a trivial classifier can achieve high accuracy.

How to eliminate wrong answers

Option A is wrong because Mean Squared Error (MSE) is a regression metric that measures average squared differences between predicted and actual values, not suitable for binary classification tasks like loan default prediction. Option C is wrong because accuracy is misleading on imbalanced datasets; a model predicting all non-default would achieve 95% accuracy but fail to identify any actual defaults. Option D is wrong because R-squared is a regression metric that indicates the proportion of variance explained by the model, inappropriate for evaluating classification performance on imbalanced data.

944
MCQmedium

A financial institution is deploying an AI model that predicts loan default risk. The model is trained on historical data that includes sensitive attributes like zip code and marital status. The compliance team is concerned about disparate impact. Which technique should be applied during model training to mitigate bias while maintaining predictive performance?

A.Apply adversarial debiasing during training to minimize the model's ability to predict the sensitive attribute.
B.Post-process predictions by adjusting thresholds for different demographic groups.
C.Remove all sensitive attributes from the training dataset.
D.Oversample the minority group in the training data to balance representation.
AnswerA

Adversarial debiasing trains the main model to make accurate predictions while simultaneously training an adversary to predict the sensitive attribute from the model's outputs. By minimizing the adversary's success, the model learns representations that are less informative about sensitive attributes, reducing disparate impact. This technique can be integrated into training and often preserves predictive performance better than simply removing sensitive features.

Why this answer

Adversarial debiasing is a training-time technique that reduces the model's reliance on sensitive attributes by pitting the predictor against an adversary. This mitigates disparate impact while often maintaining accuracy. Removing sensitive attributes is insufficient due to proxies, oversampling addresses class imbalance rather than bias, and post-processing occurs after training, which does not meet the scenario's requirement for an in-training method.

Exam trap

The trap here is assuming that removing sensitive attributes guarantees fairness, when proxy variables can perpetuate bias and more robust training-time techniques are needed.

945
MCQeasy

An organization is deploying a machine learning model that classifies loan applications. They want to prevent an attacker from reconstructing individual customer records from the model's predictions. Which type of attack should they defend against?

A.Membership inference
B.Data poisoning
C.Model inversion
D.Adversarial example
AnswerC

Model inversion reconstructs training data by querying predictions, directly threatening the customer records in the loan classifier. It differs from membership inference, which only determines whether a record was used. Limiting prediction confidence and adding noise to outputs mitigates this specific reconstruction risk.

Why this answer

Model inversion attacks allow an attacker to reconstruct the original training data by analyzing the model's predictions. In this scenario, the attacker could use the model's outputs to infer sensitive details about individual loan applicants, such as income or credit history, violating privacy. Defending against model inversion is critical when predictions can be used to reverse-engineer private training records.

Exam trap

CompTIA often tests the distinction between model inversion (reconstructing data) and membership inference (detecting presence of data), so the trap here is confusing the goal of reconstructing records with simply inferring membership.

How to eliminate wrong answers

Option A is wrong because membership inference attacks aim to determine whether a specific record was part of the training dataset, not to reconstruct the actual data values. Option B is wrong because data poisoning attacks involve corrupting the training data to manipulate model behavior, not extracting or reconstructing existing records. Option D is wrong because adversarial example attacks craft malicious inputs to cause misclassification, not to reconstruct training data from predictions.

946
MCQhard

A data scientist is training a model to detect fraudulent transactions. To protect customer privacy, the team wants to ensure that the model does not inadvertently memorize and reveal sensitive information about individuals in the training set. Which technique should be applied during training?

A.Differential privacy
B.Federated learning
C.Homomorphic encryption
D.Model quantization
AnswerA

Differential privacy adds calibrated noise to training gradients or outputs, mathematically bounding any single individual's influence on the model. This directly satisfies the stem's requirement that the model cannot memorise and reveal sensitive customer details, providing a provable privacy guarantee rather than relying on heuristic anonymisation of the transaction data.

Why this answer

Differential privacy is the correct technique because it adds calibrated noise to the training process or output, ensuring that the model cannot infer whether any specific individual's data was included in the training set. This directly addresses the goal of preventing memorization and leakage of sensitive information while still allowing the model to learn useful patterns for fraud detection.

Exam trap

The AI0-001 exam often tests the misconception that federated learning alone guarantees privacy, when in fact it only addresses data locality and must be combined with differential privacy to prevent model inversion or membership inference attacks.

How to eliminate wrong answers

Option B (Federated learning) is wrong because it focuses on training models across decentralized data without sharing raw data, but it does not inherently prevent the model from memorizing individual records; additional privacy techniques like differential privacy are needed. Option C (Homomorphic encryption) is wrong because it enables computation on encrypted data, protecting data in transit or at rest, but it does not address model memorization or inference of training data from the model's outputs. Option D (Model quantization) is wrong because it reduces the precision of model weights to improve efficiency, but it has no effect on privacy or preventing memorization of sensitive information.

947
MCQhard

An organization wants to implement an AI ethics board. Which composition best ensures independence and expertise?

A.All members from the legal department
B.IT department head and data scientists
C.Mix of internal stakeholders and external ethicists
D.Only senior executives from the company
AnswerC

Blending internal stakeholders with external ethicists satisfies the independence constraint: external members lack reporting lines to the organisation, so they can challenge internal priorities without career risk, while internal stakeholders supply operational context. This dual composition delivers both the detached scrutiny and domain expertise an ethics board requires.

Why this answer

An AI ethics board must combine internal stakeholders (who understand organizational context, data flows, and operational constraints) with external ethicists (who provide independent, unbiased perspectives and specialized knowledge of ethical frameworks like IEEE Ethically Aligned Design or the EU AI Act). This composition ensures the board can evaluate AI systems for bias, fairness, and transparency without being dominated by business or technical interests, which is critical for maintaining trust and regulatory compliance.

Exam trap

The AI0-001 exam often tests the misconception that technical expertise alone (Option B) is sufficient for AI ethics governance, but the trap is that independence and multidisciplinary perspectives are explicitly required to avoid conflicts of interest and ensure comprehensive ethical evaluation.

How to eliminate wrong answers

Option A is wrong because a board composed solely of legal department members lacks the technical expertise to assess AI model behavior, data provenance, and algorithmic bias, and may focus narrowly on legal compliance rather than broader ethical principles. Option B is wrong because IT department heads and data scientists bring deep technical knowledge but have inherent conflicts of interest (e.g., pressure to deploy models quickly) and lack the independent ethical oversight needed to challenge internal decisions. Option D is wrong because senior executives prioritize business outcomes and shareholder value, which can compromise impartiality and lead to ethics being subordinated to profit, violating the independence required for effective governance.

948
MCQmedium

A retail company uses a gradient boosting model to predict customer lifetime value (CLV). The model currently uses 50 features including purchase history, demographics, and web behavior. The model's RMSE on the test set is 120. The data science team wants to improve the model's accuracy without increasing training time significantly. They have access to additional data: customer support interaction logs (text), social media sentiment (text), and third-party credit scores (numeric). They also have the ability to perform feature engineering, hyperparameter tuning, and ensemble methods. Which approach is most likely to yield the best improvement in predictive performance with minimal increase in training time?

A.Add the customer support text as a feature using TF-IDF vectors
B.Use an ensemble of gradient boosting and random forest models
C.Perform hyperparameter tuning using grid search
D.Engineer new features such as average purchase value and recency
AnswerD

Feature engineering can capture patterns without adding new data sources or significant time.

Why this answer

Engineering domain-relevant features like average purchase value and recency directly captures the underlying behavioral patterns that drive customer lifetime value, often providing a higher signal-to-noise ratio than adding raw text or third-party data. This approach leverages existing data without significantly increasing the feature dimensionality or training time, unlike adding TF-IDF vectors which would dramatically expand the feature space and slow training.

Exam trap

CompTIA often tests the misconception that adding more data (especially text) or complex ensemble methods always improves model accuracy, while the correct approach is to engineer features that capture domain-specific patterns with minimal computational overhead.

How to eliminate wrong answers

Option A is wrong because adding customer support text as TF-IDF vectors would introduce thousands of sparse features, significantly increasing training time and risking overfitting without guaranteed improvement in RMSE. Option B is wrong because ensembling gradient boosting with random forest typically increases training time substantially (both models must be trained) and may not outperform a well-tuned single gradient boosting model on structured data. Option C is wrong because hyperparameter tuning using grid search is computationally expensive, often requiring many model fits, and would increase training time more than feature engineering without leveraging the new data sources.

949
Multi-Selecthard

A company uses an AI model to screen job applicants. A disparate impact analysis reveals that the model's rejection rate for a protected group is significantly higher than for others. Which THREE actions should the company take to address this?

Select 3 answers
A.Revisit training data for historical bias and consider reweighting
B.Ignore the disparity because the model is accurate overall
C.Remove all demographic attributes from the dataset
D.Apply fairness constraints or adversarial debiasing during training
E.Consider using a different model that achieves better fairness metrics
AnswersA, D, E

Addressing data bias is a fundamental step to reduce disparate impact.

Why this answer

Revisiting the training data for historical bias and applying reweighting directly addresses the root cause of disparate impact. If the training data contains biased labels or skewed representation of the protected group, the model will learn and amplify those biases. Reweighting adjusts the loss function to give more importance to underrepresented or disadvantaged groups, helping to equalize error rates across groups.

Exam trap

The AI0-001 exam often tests the misconception that removing protected attributes (option C) is sufficient to eliminate bias, when in reality it can hide bias and still allow proxy discrimination, making it an incomplete and sometimes counterproductive solution.

950
MCQhard

A data scientist notices the model overfits. Which change to the exhibit's configuration would most likely reduce overfitting?

A.Remove dropout layers
B.Increase learning rate to 0.01
C.Add L2 regularization to dense layers
D.Increase units in the first dense layer to 512
AnswerC

Adding L2 regularization to dense layers penalises large weights in the loss function, reducing the network's capacity to fit training noise. This directly counteracts the overfitting observed in the exhibit's configuration without altering the architecture.

Why this answer

Adding L2 regularization to dense layers penalizes large weights by adding a squared magnitude term to the loss function, which forces the model to learn simpler patterns and reduces overfitting. This directly addresses the core issue of the model memorizing noise in the training data.

Exam trap

CompTIA often tests the misconception that increasing model capacity (more units or layers) or removing regularization always improves performance, when in fact these changes exacerbate overfitting; candidates must recognize that regularization techniques like L2 are specifically designed to penalize complexity and reduce overfitting.

How to eliminate wrong answers

Option A is wrong because removing dropout layers would actually increase overfitting, as dropout is a regularization technique that randomly drops neurons during training to prevent co-adaptation. Option B is wrong because increasing the learning rate to 0.01 (a relatively high value) can cause the optimizer to overshoot minima and lead to unstable training, but it does not directly reduce overfitting; in fact, a too-high learning rate may prevent convergence altogether. Option D is wrong because increasing units in the first dense layer to 512 adds more parameters to the model, which increases capacity and typically worsens overfitting rather than reducing it.

951
MCQmedium

A team is training a image classification model. They split the dataset into training, validation, and test sets. After training, the model achieves 98% accuracy on the training set but only 72% on the test set. Which step in the AI project lifecycle should the team focus on?

A.Data acquisition – collect more data
B.Model selection – use regularization or reduce model complexity
C.Deployment – re-deploy with a different serving framework
D.Data preparation – check for train/test leakage
AnswerB

Regularisation or reduced complexity directly addresses the 26-point train–test gap, which signals overfitting: the model has memorised training noise rather than learning generalisable features. Penalising large weights or shrinking capacity lowers training accuracy while raising test accuracy, satisfying the stem's requirement to close that generalisation gap.

Why this answer

The 98% training accuracy versus 72% test accuracy is a classic signature of overfitting — the model has memorized the training data rather than learning generalizable patterns. The appropriate lifecycle response is to address model complexity through regularization (L1/L2, dropout, early stopping) or by reducing the number of parameters/layers. This directly targets the generalization gap rather than the data pipeline or deployment.

Exam trap

AI0-001 often tests the difference between overfitting (high train, lower test) and underfitting (low train, low test), and candidates frequently jump to 'collect more data' when the real fix is regularization or reduced model complexity.

How to eliminate wrong answers

Option A is wrong because collecting more data would not fix overfitting if the model is already memorizing the training set — the gap is caused by model capacity, not data volume. Option C is wrong because deployment/serving frameworks have no effect on the train-test accuracy gap; the model's learned weights are the problem, not how they are served. Option D is wrong because train/test leakage would typically inflate test accuracy (making it suspiciously high), not produce a large train-test gap where training is much higher than test.

952
MCQeasy

Under the GDPR, individuals have the right to not be subject to a decision based solely on automated processing if it produces legal effects. Which of the following is a typical safeguard that organisations must provide to comply with this right?

A.The right to have all personal data deleted immediately
B.The right to receive a detailed mathematical explanation of the model
C.The right to demand a more favourable automated decision
D.The right to obtain human intervention on the part of the controller
AnswerD

Human intervention lets a person review and override an automated decision, directly satisfying the GDPR safeguard against solely automated processing producing legal effects. It restores meaningful human involvement, which the regulation requires alongside rights to contest and express a viewpoint.

Why this answer

Option D is correct because GDPR Article 22(3) explicitly provides that where solely automated decision-making (including profiling) produces legal or similarly significant effects, the data subject has the right to obtain human intervention, express their point of view, and contest the decision. Human intervention is the canonical safeguard organisations must offer.

Exam trap

AI0-001 often tests the confusion between Article 22 safeguards (human intervention, contest) and other GDPR rights (erasure, access, explanation) — candidates who pick 'right to erasure' or 'mathematical explanation' conflate distinct articles and recitals.

How to eliminate wrong answers

Option A is wrong because the right to erasure (Article 17) is a separate right and is not the safeguard tied to Article 22 — and it is not absolute (it has exceptions for legal obligations, public interest, etc.). Option B is wrong because GDPR Recital 71 refers to 'meaningful information about the logic involved', not a detailed mathematical explanation of the model — trade secrets and the impracticality of explaining deep models mean a full mathematical explanation is not required. Option C is wrong because GDPR does not grant a right to a 'more favourable' automated decision — it grants the right to contest and to human intervention, not to a preferred outcome.

953
MCQeasy

A startup is developing an AI chatbot and wants to use a pre-trained language model to generate responses. They need to integrate the model into their application with minimal latency and cost. Which approach should they take?

A.Use a rule-based chatbot system
B.Train a custom transformer model from scratch on their own data
C.Use a pre-trained model via a managed API service
D.Deploy the model on a local server with a single GPU
AnswerC

Using a pre-trained model through a managed API service, such as OpenAI's GPT or Google's Vertex AI, allows the startup to integrate advanced language capabilities without managing infrastructure. This approach minimizes latency by leveraging the provider's optimized serving and reduces cost by paying only for usage, making it ideal for a startup.

Why this answer

The startup needs to integrate a pre-trained language model with minimal latency and cost. Using a managed API service provides immediate access to state-of-the-art models without infrastructure management, and the pay-as-you-go model keeps costs low. This is the most efficient approach for a startup with limited resources.

Exam trap

The trap here is overlooking the operational and cost benefits of managed API services and instead opting for self-hosted solutions that require significant upfront investment.

954
MCQmedium

A city agency deploys an AI system that scores permit applications. The vendor refuses to disclose model weights or feature importance, citing trade secrets. The agency's oversight board must still meet its obligation to explain adverse decisions to applicants. Which approach best satisfies that obligation?

A.Require the vendor to supply model-agnostic explanations, such as local surrogate or counterfactual reason codes, for each adverse decision.
B.Inform applicants that the decision was made by an automated system and provide a generic appeal link.
C.Replace the vendor model with a transparent rule-based scoring system that the agency builds in-house.
D.Publish the vendor's source code and trained weights so independent researchers can audit the decisions.
AnswerA

Model-agnostic techniques generate explanations from input-output behavior without exposing proprietary weights, so the agency can give applicants concrete reasons while the vendor keeps its intellectual property. Local surrogates and counterfactual reason codes translate a decision into interpretable factors. This satisfies the oversight duty and preserves the commercial relationship, which is exactly the constraint the scenario describes.

Why this answer

The agency must explain adverse decisions without forcing the vendor to reveal proprietary internals. Model-agnostic explainability, including local surrogates and counterfactual reason codes, derives per-decision explanations from observable behavior, so applicants receive meaningful factors and the board meets its duty. Full disclosure, generic notices, and replacing the system each either breach the constraint or fail to provide decision-specific reasoning.

Exam trap

The trap here is believing that meaningful explanation requires access to model weights, when model-agnostic techniques can produce per-decision reasons without disclosing proprietary internals.

955
MCQeasy

A team is using a pre-trained language model for sentiment analysis. They want to adapt it to a specific domain with limited labeled data. Which approach is most efficient?

A.Fine-tune the pre-trained model on domain data
B.Use the pre-trained model as is
C.Train a new model from scratch
D.Ensemble multiple pre-trained models
AnswerA

Fine-tuning updates the pre-trained weights on the small domain-labelled set, adapting the model's representations to domain-specific vocabulary and sentiment patterns. This leverages the existing general language knowledge, so far fewer labelled examples are needed than training from scratch.

Why this answer

Fine-tuning a pre-trained language model on domain-specific labeled data is the most efficient approach because it leverages the general language understanding learned from large corpora while adapting to the target domain with minimal additional data. This process uses transfer learning, where only the final layers or a subset of parameters are updated, significantly reducing the amount of labeled data and compute required compared to training from scratch.

Exam trap

The AI0-001 exam often tests the misconception that a pre-trained model can be used directly for any domain without adaptation, leading candidates to choose Option B, but the trap here is that domain-specific tasks require fine-tuning to align the model's representations with the target data distribution.

How to eliminate wrong answers

Option B is wrong because using the pre-trained model as-is (zero-shot inference) typically yields poor performance on domain-specific sentiment analysis due to vocabulary and context mismatches, as the model was not exposed to domain-specific jargon or sentiment nuances. Option C is wrong because training a new model from scratch requires a massive labeled dataset (often millions of examples) and extensive computational resources, which contradicts the constraint of limited labeled data. Option D is wrong because ensembling multiple pre-trained models without fine-tuning them on domain data does not address the domain adaptation problem; it merely averages their general predictions, which may still be inaccurate for domain-specific sentiment.

956
MCQeasy

A company has developed a deep learning model for image classification. The team wants to deploy the model to production with high availability and scalability. Which approach should they use?

A.Run the model on a laptop during business hours.
B.Deploy the model as a monolithic application on a single server.
C.Embed the model directly into a mobile app.
D.Use a containerized approach with Kubernetes.
AnswerD

Containerising the model and orchestrating it with Kubernetes delivers horizontal pod autoscaling and self-healing replicas, directly satisfying the stem's high-availability and scalability requirements. Unlike a single VM or serverless endpoint, Kubernetes spreads inference pods across nodes, so demand spikes and node failures do not interrupt serving.

Why this answer

Containerization with Kubernetes provides the orchestration, auto-scaling, and self-healing capabilities required for high availability and scalability in production. Kubernetes manages container lifecycles, distributes traffic across replicas via Services and Ingress controllers, and can automatically scale pods based on CPU/memory metrics or custom metrics, ensuring the deep learning model handles variable loads without downtime.

Exam trap

CompTIA often tests the misconception that embedding AI models directly into mobile apps or running them on a single server is sufficient for production, when in reality enterprise-grade deployments require container orchestration for resilience and elasticity.

How to eliminate wrong answers

Option A is wrong because running the model on a laptop during business hours lacks any production-grade availability, scalability, or fault tolerance; it is a single point of failure and cannot handle concurrent requests. Option B is wrong because a monolithic application on a single server creates a single point of failure, cannot scale horizontally, and offers no load balancing or automated recovery, making it unsuitable for high availability. Option C is wrong because embedding the model directly into a mobile app offloads inference to client devices, which introduces latency, security risks, and inconsistent performance; it does not provide centralized high availability or scalability for the production service.

957
MCQhard

A machine learning engineer is training a deep neural network for image classification. The training loss decreases steadily, but the validation loss starts to increase after 20 epochs. The engineer wants to implement a technique that dynamically adjusts the learning rate during training to improve convergence and generalization. Which method should the engineer use?

A.Stochastic Gradient Descent (SGD) with a fixed learning rate
B.Batch Normalization
C.Dropout regularization
D.Learning Rate Scheduler with exponential decay
AnswerD

An exponential decay learning rate scheduler reduces the learning rate over time, which can help the model settle into a deeper minimum and improve generalization. As training progresses, smaller learning rates allow finer adjustments to weights, potentially mitigating the increasing validation loss by preventing overshooting and encouraging convergence to a smoother minimum.

Why this answer

A learning rate scheduler with exponential decay dynamically reduces the learning rate as training progresses, which helps the model converge more smoothly and can improve generalization. This directly addresses the need to adjust the learning rate during training to combat the rising validation loss, unlike fixed learning rates or other regularization techniques that do not modify the learning rate.

Exam trap

The trap here is confusing regularization techniques like dropout with optimization techniques that adjust the learning rate, even though both can help with overfitting.

958
MCQeasy

Which neural network architecture is specifically designed to handle sequential data and mitigate the vanishing gradient problem?

A.Convolutional Neural Network (CNN)
B.Transformer
C.Vanilla Recurrent Neural Network (RNN)
D.Long Short-Term Memory (LSTM) network
AnswerD

LSTM networks use gated cells (input, forget, output) that regulate information flow across time steps, preserving gradients over long sequences. This architecture directly mitigates the vanishing gradient problem while handling sequential data, matching the stem's requirement.

Why this answer

LSTM (Long Short-Term Memory) is a type of RNN designed with gating mechanisms to prevent vanishing gradients in long sequences. CNNs are for spatial data; vanilla RNNs suffer from vanishing gradients; transformers use attention but are not specifically designed to mitigate vanishing gradients (they use residual connections).

959
Multi-Selecthard

A company is deploying an AI-based resume screening tool. The security team is concerned about adversarial attacks that could manipulate the tool's rankings. Which TWO of the following are effective defenses against such attacks? (Choose two.)

Select 2 answers
A.Input sanitization to remove special characters
B.Model extraction prevention via API rate limiting
C.Differential privacy during training
D.Certified robustness via randomized smoothing
E.Adversarial training with perturbed resumes
AnswersD, E

Randomized smoothing is a certified defense that provides provable robustness guarantees against small input perturbations. For resume screening, it could ensure that minor changes to a resume do not drastically alter the ranking. This technique is effective against adversarial attacks and adds a layer of security.

Why this answer

Adversarial training and certified robustness via randomized smoothing are both proactive defenses that make the model more resistant to crafted inputs. Adversarial training exposes the model to perturbed examples during training, while randomized smoothing provides a formal guarantee against small perturbations. These methods directly counter evasion attacks that could manipulate resume rankings.

Exam trap

The trap here is confusing privacy-preserving techniques like differential privacy with defenses against adversarial manipulation.

960
MCQmedium

An AI team notices that their hiring model consistently selects male candidates over equally qualified female candidates. Analysis shows the training data contains past hiring decisions where men were predominantly hired. Which type of bias is the root cause?

A.Algorithmic bias
B.Confirmation bias
C.Selection bias
D.Historical bias
AnswerD

Historical bias arises when training data reflects past discriminatory decisions, so the model learns and reproduces those patterns. Here the data encodes prior hiring favouring men, which the model replicates against equally qualified women. This directly satisfies the stem's constraint: bias originates in the data itself, not the algorithm or deployment.

Why this answer

Historical bias is the root cause because the training data reflects past hiring decisions that systematically favored male candidates, encoding societal or organizational prejudices into the model. The model learns these historical patterns and perpetuates them, leading to discriminatory outcomes against equally qualified female candidates. This is distinct from algorithmic bias, which would arise from the model's design or optimization process itself.

Exam trap

CompTIA AI often tests the distinction between historical bias (data-driven) and algorithmic bias (model-driven), and the trap here is that candidates may confuse the source of bias as being from the algorithm itself rather than the training data.

How to eliminate wrong answers

Option A is wrong because algorithmic bias refers to bias introduced by the algorithm's design, training process, or optimization function, not by the data itself. Option B is wrong because confirmation bias is a cognitive bias where individuals favor information that confirms their preexisting beliefs, which is not applicable to a machine learning model's training data. Option C is wrong because selection bias occurs when the data is not representative of the population due to non-random sampling, but here the data accurately reflects historical hiring decisions, which are themselves biased.

961
MCQhard

A computer vision engineer is building a model to detect defects on a manufacturing line. Defects are rare, occurring in only 0.5% of images. The engineer trains a convolutional neural network and achieves 99.5% accuracy, but the model never predicts a defect. The engineer wants to address the underlying issue. Which approach is MOST appropriate?

A.Switch from accuracy to mean squared error as the evaluation metric
B.Increase the learning rate to help the model escape the majority-class solution
C.Apply class weighting or resampling techniques to emphasize the minority defect class
D.Add more convolutional layers to increase model capacity
AnswerC

Class weighting or resampling directly counteracts the imbalance by increasing the cost of misclassifying defects or by presenting more defect examples during training. This forces the model to learn defect features instead of defaulting to the majority class. It is the standard, targeted remedy for a model that achieves high accuracy but fails on the rare class.

Why this answer

The model achieves high accuracy by always predicting the majority class, a classic symptom of severe class imbalance. Class weighting or resampling techniques adjust the training process so that minority-class errors carry more weight or appear more frequently, compelling the model to learn defect patterns. The other options do not target the imbalance and are unlikely to resolve the failure to detect rare defects.

Exam trap

The trap here is treating high accuracy as evidence of a good model when the metric is misleading under class imbalance, leading to fixes that ignore the need to rebalance the training signal.

962
Multi-Selecthard

A natural language processing team is building a sentiment analysis model for customer reviews. They want to ensure the model generalizes well to new, unseen reviews and does not simply memorize the training data. Which TWO techniques are most appropriate to achieve this goal? (Choose two.)

Select 2 answers
A.Remove all stop words from the reviews before training.
B.Use a larger batch size during training.
C.Apply L2 regularization to the model's weights.
D.Increase the number of training epochs until training loss approaches zero.
E.Implement early stopping based on validation loss.
AnswersC, E

L2 regularization adds a penalty proportional to the square of the weights to the loss function, discouraging large weights and reducing model complexity. This helps prevent overfitting by forcing the model to learn smoother, more generalizable patterns rather than memorizing training examples. For sentiment analysis, it is a standard and effective technique to improve performance on unseen reviews.

Why this answer

L2 regularization penalizes large weights, encouraging simpler models that generalize better, while early stopping halts training when validation performance degrades, preventing overfitting. Both directly target the goal of avoiding memorization of training data. Increasing epochs, larger batch sizes, and stop word removal do not reliably improve generalization and may even hurt performance.

Exam trap

The trap here is assuming that any technique that speeds up training or reduces data size will also improve generalization, when in fact only methods that constrain model complexity or monitor validation performance directly address overfitting.

Page 12

Page 13 of 13