Courseiva

CompTIA AI+ AI0-001 (AI0-001) — Questions 76–150

962 questions total · 13pages · All types, answers revealed

Page 1

Page 2 of 13

Page 3
76
MCQhard

An AI system used for autonomous driving is found to have a lower accuracy in detecting pedestrians with darker skin tones. The development team wants to address this ethical issue. Which action is most effective?

A.Conduct additional testing to measure the disparity
B.Augment the training dataset with more images of pedestrians with darker skin
C.Replace the object detection algorithm with a different one
D.Adjust the model's decision threshold for pedestrian detection
AnswerB

Adding more darker-skinned pedestrian images directly corrects the class imbalance in the training distribution, which is the root cause of the disparate accuracy. The model learns underrepresented features only when sufficient examples exist, so dataset augmentation reduces the bias more effectively than post-hoc threshold tuning or documentation.

Why this answer

Augmenting the training dataset with more images of pedestrians with darker skin directly addresses the root cause of the bias: underrepresentation in the training data. By providing a more balanced and diverse dataset, the model can learn more robust features for all skin tones, reducing accuracy disparity without altering the algorithm's core logic or introducing arbitrary thresholds.

Exam trap

CompTIA often tests the misconception that bias can be fixed by simply changing the algorithm or threshold, when in reality the most effective first step is to address data imbalance through targeted augmentation.

How to eliminate wrong answers

Option A is wrong because additional testing only measures the disparity but does not fix it; it is a diagnostic step, not a corrective action. Option C is wrong because replacing the object detection algorithm does not guarantee improved fairness—bias often stems from training data distribution, not the algorithm itself, and a different algorithm may still exhibit similar biases if trained on the same skewed data. Option D is wrong because adjusting the decision threshold can trade off precision and recall but does not address the underlying data imbalance; it may reduce false negatives for one group at the expense of increased false positives for another, without resolving the root cause.

77
Multi-Selecthard

A company is integrating a third-party pre-trained model into its product. To address supply chain security, which THREE actions are most important? (Choose three.)

Select 3 answers
A.Checking the model for backdoors using validation techniques
B.Using homomorphic encryption for model inference
C.Creating a software bill of materials (SBOM) for AI components
D.Implementing federated learning for future updates
E.Vetting the model's provenance and dataset lineage
AnswersA, C, E

Validating the pre-trained model for backdoors detects trojaned weights or triggers that activate malicious behaviour after integration. This addresses supply chain security by verifying the third-party artefact before deployment, since provenance alone cannot guarantee the model is free of implanted malicious functionality.

Why this answer

Option A is correct because validating a third-party pre-trained model for backdoors (e.g., via trigger-pattern scanning, anomaly detection, or red-team testing) directly mitigates the risk that a maliciously tampered model contains hidden behaviors that activate on specific inputs. Option C is correct because an SBOM for AI components enumerates the model's dependencies, libraries, weights, and versions, giving the organization the transparency needed to track and remediate vulnerabilities across the supply chain. Option E is correct because vetting the model's provenance and dataset lineage verifies where the model and its training data came from, ensuring they originate from trusted sources and have not been poisoned or tampered with.

Option B is not appropriate here because homomorphic encryption protects data during inference but does not address supply chain integrity of the model itself. Option D is also not appropriate because federated learning is a training architecture for future updates and does not secure the initial integration of a third-party pre-trained model.

Exam trap

CompTIA often tests the distinction between supply chain security (provenance, SBOM, backdoor checks) and operational security (encryption, federated learning), so candidates mistakenly pick options that sound security-related but address different threat models.

78
MCQhard

A data scientist is training a random forest model on a large dataset and notices that the model is overfitting. Which hyperparameter adjustment is most likely to reduce overfitting?

A.Increase the maximum features
B.Decrease the maximum depth of trees
C.Decrease the minimum samples split
D.Increase the number of trees
AnswerB

Maximum depth controls how many splits each tree can make. Reducing it constrains tree complexity, limiting the model's ability to memorise training noise, which lowers variance and reduces overfitting. Increasing depth or adding trees would typically worsen the overfitting observed.

Why this answer

Decreasing the maximum depth of trees limits how deep each decision tree can grow, which reduces the model's capacity to learn overly specific patterns from the training data. This directly combats overfitting by enforcing simpler trees that generalize better to unseen data.

Exam trap

A common mistake is assuming that increasing the number of trees always reduces overfitting, but in random forests, while more trees reduce variance through averaging, they do not address the root cause of overfitting in individual trees. Decreasing maximum depth directly limits tree complexity.

How to eliminate wrong answers

Option A is wrong because increasing the maximum features (the number of features considered at each split) actually increases tree diversity and can reduce overfitting in some cases, but it is not the most direct adjustment for overfitting—it can also increase variance if set too high. Option C is wrong because decreasing the minimum samples split (the minimum number of samples required to split an internal node) allows trees to split on smaller subsets, which increases model complexity and exacerbates overfitting. Option D is wrong because increasing the number of trees generally improves stability and reduces variance due to averaging, but it does not directly reduce overfitting; in fact, more trees can sometimes memorize noise if individual trees are already overfit.

79
MCQhard

An AI governance committee is reviewing a resume-screening model. The model's accuracy is high overall, but its false negative rate is much higher for applicants from one demographic group than for others. The committee wants to address this disparity. Which action best targets the problem?

A.Increase the model's overall accuracy by adding more training data
B.Remove all demographic attributes from the training data
C.Measure and mitigate bias across demographic groups, including error rates and representation
D.Replace the model with a simpler algorithm that is inherently interpretable
AnswerC

Disparate false negative rates across groups are a fairness and bias signal. Auditing error rates by group, checking training data representation, and applying mitigation such as reweighting or threshold adjustment directly addresses the observed disparity, whereas overall accuracy can hide unequal performance.

Why this answer

The scenario describes unequal false negative rates, a fairness problem that requires measuring performance by demographic group and mitigating the disparity through data or threshold adjustments. Adding data, removing protected attributes, or switching to a simpler algorithm may change accuracy or transparency but does not directly close the group-level error gap.

Exam trap

The trap here is believing that deleting protected attributes makes a model fair, when proxy variables preserve the bias and hide it from measurement.

80
Multi-Selectmedium

A team is selecting a vector database for a RAG application that requires low-latency similarity search on millions of embeddings. They prioritize ease of use and fully managed cloud service. Which TWO options meet these requirements?

Select 2 answers
A.Pinecone
B.pgvector
C.Chroma
D.Weaviate
E.Milvus
AnswersA, D

Pinecone is a fully managed, cloud-native vector database with low-latency similarity search, ideal for production RAG.

Why this answer

Pinecone is a fully managed vector database designed for production-scale RAG applications, offering low-latency similarity search on millions of embeddings without requiring users to manage infrastructure. Its serverless architecture and simple API align directly with the team's priorities of ease of use and a fully managed cloud service. Weaviate also offers a fully managed cloud service (Weaviate Cloud) with low-latency similarity search and an easy-to-use API, making it a second valid choice.

By contrast, pgvector is a PostgreSQL extension that is not itself a fully managed cloud service; Chroma is primarily an open-source embedded/local database; and Milvus is typically self-managed (Zilliz Cloud is a separate managed offering), so these do not inherently satisfy the 'fully managed cloud service' requirement.

Exam trap

CompTIA often tests the distinction between open-source, self-managed tools and fully managed cloud services, where candidates may incorrectly assume that any popular vector database (like Milvus or pgvector) inherently provides a managed cloud experience without checking the deployment model.

81
MCQeasy

What is the primary function of an AI ethics board within an organization?

A.Developing algorithms
B.Managing cloud infrastructure
C.Marketing AI products
D.Reviewing AI projects for ethical compliance
AnswerD

An AI ethics board's core remit is governance: examining proposed and live AI projects against ethical principles, flagging risks such as bias or privacy harm, and advising on remediation. It reviews for compliance rather than building models or owning day-to-day operations.

Why this answer

The primary function of an AI ethics board is to review AI projects for ethical compliance, ensuring that the organization's AI systems adhere to established ethical principles, legal standards, and governance frameworks. This board typically assesses risks related to bias, fairness, transparency, and accountability before deployment, rather than engaging in technical development or operational tasks.

Exam trap

The AI0-001 exam often tests the distinction between operational roles (e.g., development, infrastructure, marketing) and governance roles (e.g., ethics review), so the trap here is confusing a technical or business function with the oversight responsibility of an ethics board.

How to eliminate wrong answers

Option A is wrong because developing algorithms is a technical function performed by data scientists and engineers, not by an ethics board, which focuses on governance and oversight. Option B is wrong because managing cloud infrastructure is an IT operations role involving platforms like AWS or Azure, unrelated to ethical review processes. Option C is wrong because marketing AI products is a business development activity that promotes AI solutions, whereas an ethics board provides independent scrutiny to prevent unethical practices.

82
MCQhard

A healthcare AI team is training a model to predict patient readmission risk from electronic health records. The dataset contains sensitive patient data and must comply with HIPAA. They need to ensure that the model training process does not expose protected health information (PHI) and that the model does not memorize individual patient data. Which technique should they implement?

A.Data augmentation with synthetic records
B.Federated learning across hospitals
C.Homomorphic encryption of the model weights
D.Differential privacy during training
AnswerD

Differential privacy adds controlled noise to the training process, ensuring that the inclusion or exclusion of any single patient's data does not significantly affect the model's output. This prevents memorization of individual records and helps comply with HIPAA by protecting PHI. It is a rigorous mathematical framework for privacy-preserving machine learning.

Why this answer

Differential privacy is the only technique listed that provides a formal guarantee that the model does not memorize individual patient data. By injecting noise during training, it ensures that the model's predictions are statistically similar whether or not any single patient's data is included. This directly addresses the HIPAA compliance and privacy concerns in the scenario.

Exam trap

The trap here is confusing privacy-preserving techniques that protect data in transit or storage with those that prevent model memorization during training.

83
Multi-Selecthard

A company is building a multi-modal AI application that processes text, images, and audio. They need a unified platform to store embeddings for all modalities, perform hybrid search (vector + metadata filtering), and scale to millions of vectors. Which THREE services are suitable for this purpose? (Choose THREE.)

Select 3 answers
A.Weaviate
B.Amazon S3
C.Snowflake
D.pgvector (PostgreSQL extension)
E.Pinecone
AnswersA, D, E

Weaviate is a vector database with hybrid search and multi-modal support.

Why this answer

Weaviate is a purpose-built vector database that natively supports multi-modal embeddings (text, images, audio) through its vectorizer modules and provides hybrid search combining vector similarity with metadata filtering (e.g., using GraphQL or REST APIs). It is designed to scale to millions of vectors with built-in sharding and replication, making it a strong fit for the described unified platform.

Exam trap

CompTIA often tests the distinction between general-purpose storage (S3) or analytics platforms (Snowflake) and purpose-built vector databases, leading candidates to mistakenly choose services that store data but lack native vector search and hybrid filtering capabilities.

84
Multi-Selectmedium

A developer is building an AI agent that needs to call external tools (e.g., weather API, database) and reason about the results to answer user queries. Which THREE components are essential for implementing this agentic workflow?

Select 3 answers
A.Planning capability (e.g., step-by-step decomposition)
B.ReAct (Reasoning + Acting) loop
C.Fine-tuned domain-specific model
D.A vector store for long-term memory
E.Function calling or tool use interface
AnswersA, B, E

Planning capability decomposes a multi-step query into ordered sub-tasks, letting the agent decide which external tool to invoke and in what sequence. Without step-by-step decomposition, the agent cannot coordinate the weather API and database calls needed to reason over combined results before answering.

Why this answer

Option A (Planning capability) is correct because an agentic workflow requires the model to decompose a complex user query into an ordered sequence of steps, deciding which sub-tasks to execute and in what order before invoking tools. Option B (ReAct loop) is correct because the Reasoning + Acting pattern interleaves thought, action, and observation cycles, letting the agent call a tool, reason over the returned result, and decide the next action until the query is resolved. Option E (Function calling or tool use interface) is correct because the agent must have a structured mechanism (e.g., JSON schema-based function/tool definitions) to invoke the weather API or database and receive machine-readable outputs.

Option C (Fine-tuned domain-specific model) is not essential, since a general-purpose LLM with tool-calling and reasoning prompts can drive the workflow without domain fine-tuning. Option D (Vector store for long-term memory) is not essential either, as retrieval-based long-term memory is an optional enhancement rather than a required component for calling tools and reasoning over their results.

Exam trap

The AI0-001 exam often tests the misconception that fine-tuning or vector stores are mandatory for agentic workflows, when in fact the core requirements are planning, a reasoning-acting loop, and a tool-use interface, all achievable with a base model and prompt engineering.

85
MCQhard

A company is building a computer vision system to detect defects in manufactured parts. They have 10,000 labeled images per class (defective and non-defective). They want to achieve high accuracy with limited computational resources. Which deep learning architecture and approach is most appropriate?

A.Train a custom CNN from scratch with many layers
B.Use a decision tree ensemble
C.Use a pre-trained VGG16 and fine-tune the last few layers
D.Use an RNN to process image sequences
AnswerC

Transfer learning with a pre-trained VGG16 exploits features already learned from large image datasets, so fine-tuning only the final layers reaches high defect-detection accuracy with far less training data and compute than training a deep network from scratch.

Why this answer

Using a pre-trained VGG16 and fine-tuning the last few layers leverages transfer learning, which is ideal when you have a moderate-sized labeled dataset (10,000 images per class) and limited computational resources. The pre-trained model already captures general visual features from ImageNet, so only the task-specific layers need to be trained, reducing training time and resource requirements while still achieving high accuracy.

Exam trap

CompTIA AI often tests the misconception that more layers or training from scratch always yields better accuracy, when in reality transfer learning with a pre-trained model is the most practical choice for moderate datasets and limited compute.

How to eliminate wrong answers

Option A is wrong because training a custom CNN from scratch with many layers would require a very large dataset (typically millions of images) and extensive computational resources to converge, and with only 10,000 images per class, it risks overfitting and poor generalization. Option B is wrong because decision tree ensembles are not designed for high-dimensional image data; they lack the spatial feature extraction capabilities of CNNs and would perform poorly on raw pixel inputs. Option D is wrong because RNNs are designed for sequential data (e.g., time series, text) and are not suitable for static image classification; they would ignore spatial structure and be computationally inefficient for this task.

86
MCQmedium

A company is implementing a guardrail system for their LLM chatbot. Which of the following is an example of a guardrail?

A.Using a larger context window
B.Rejecting requests that ask for illegal advice
C.Increasing the model's temperature parameter
D.Enabling caching for frequent queries
AnswerB

Rejecting requests for illegal advice is a guardrail: an enforced policy that blocks disallowed outputs before they reach the user. It constrains the chatbot's behaviour to permitted content, which is exactly what a guardrail does.

Why this answer

A guardrail in an LLM system is a safety constraint that filters or rejects harmful inputs and outputs. Rejecting requests for illegal advice directly enforces policy compliance and prevents the model from generating prohibited content, which is the core function of a guardrail.

Exam trap

Candidates often confuse performance tuning parameters (context window, temperature, caching) with actual safety controls, leading them to mistake model configuration options for guardrail mechanisms.

How to eliminate wrong answers

Option A is wrong because using a larger context window increases the amount of text the model can process but does not enforce any safety or policy restrictions; it is a performance parameter, not a guardrail. Option C is wrong because increasing the model's temperature parameter controls randomness in output generation and has no role in blocking harmful or illegal requests; it is a generation hyperparameter, not a safety mechanism. Option D is wrong because enabling caching for frequent queries improves response latency and reduces computational load but does not filter or reject any content; it is an optimization technique, not a guardrail.

87
MCQmedium

A healthcare AI system misdiagnosed patients due to adversarial inputs. What security measure should be prioritized?

A.Encrypt all patient data
B.Use stronger authentication
C.Regular software updates
D.Implement adversarial training
AnswerD

Adversarial training augments the training set with perturbed examples, hardening the model's decision boundaries against the malicious inputs that caused misdiagnosis. This directly satisfies the healthcare scenario's requirement to withstand adversarial manipulation, unlike filtering or monitoring, which detect but do not immunise the model.

Why this answer

(Implement adversarial training) is correct because adversarial training makes the model robust to input manipulation. Option A (Encrypt all patient data) protects data privacy but not model integrity. Option B (Use stronger authentication) is for access control.

Option C (Regular software updates) is general maintenance and does not specifically address adversarial inputs.

88
MCQhard

You are a security engineer at a large e-commerce company that uses an AI-based recommendation system. The system is deployed on a Kubernetes cluster and uses a TensorFlow model served via REST API. Recently, the security team detected unusual API calls that caused the model to return incorrect recommendations. Analysis shows that the inputs were crafted to maximize prediction error. The team suspects an adversarial attack. You need to implement a solution that detects and mitigates such attacks in real-time without requiring model retraining. Which approach should you take?

A.Implement an input validation filter to detect and block anomalous inputs
B.Increase the number of model replicas to distribute the load
C.Retrain the model with adversarial examples
D.Roll back the model to a previous version that was not attacked
AnswerA

An input validation filter inspects incoming REST payloads against learned feature distributions, flagging perturbations that maximise prediction error before inference. This satisfies the real-time constraint without retraining, since detection operates on the input manifold rather than model weights. Blocking anomalous vectors prevents crafted adversarial examples from reaching the TensorFlow serving endpoint.

Why this answer

An input validation filter can detect and block adversarial inputs in real-time by analyzing statistical properties (e.g., outlier detection, perturbation magnitude) without modifying the model. This approach is lightweight, operates at the API gateway level, and does not require retraining, making it suitable for immediate deployment against crafted inputs that maximize prediction error.

Exam trap

CompTIA often tests the misconception that retraining or scaling can solve security issues, but the key constraint here is 'real-time detection without retraining,' which eliminates options that require model modification or do not address the attack vector.

How to eliminate wrong answers

Option B is wrong because increasing model replicas only distributes load and improves throughput, but does not detect or block malicious inputs; adversarial attacks exploit model vulnerabilities, not resource exhaustion. Option C is wrong because retraining with adversarial examples requires model retraining, which violates the constraint of 'without requiring model retraining' and is a longer-term solution, not real-time mitigation. Option D is wrong because rolling back to a previous version does not address the root cause; the same adversarial inputs would still be effective against the older model, and the attack vector remains unmitigated.

89
Multi-Selecthard

A retail company is deploying an AI system that generates personalized marketing copy and product recommendations. The legal team wants to align the deployment with the NIST AI Risk Management Framework's core functions. Which two activities are part of the MAP function? (Choose two.)

Select 2 answers
A.Establishing organizational policies that assign accountability for AI risk decisions across product, legal, and engineering teams.
B.Applying technical controls such as output filters and rate limits to reduce the chance of harmful generated copy reaching customers.
C.Continuously tracking key performance indicators and drift metrics after the recommendation engine enters production.
D.Categorizing the likelihood and potential impact of risks such as manipulated recommendations or exposure of inferred preferences.
E.Documenting the specific business context, affected stakeholders, and intended purpose of the recommendation engine.
AnswersD, E

MAP includes identifying and characterizing risks, including estimating likelihood and impact, so that later functions can prioritize them. Categorizing risks for the recommendation engine, such as manipulated outputs or inference of sensitive preferences, is a mapping activity. It precedes measurement and management, and it determines which risks warrant deeper analysis or controls, making it a correct MAP responsibility.

Why this answer

The MAP function establishes context and characterizes risks. Documenting business context and stakeholders, and categorizing likelihood and impact of identified risks, are core mapping activities. Applying controls belongs to MANAGE, continuous post-deployment tracking belongs to MEASURE, and assigning accountability across the organization belongs to GOVERN, which underpins all other functions.

Exam trap

The trap here is treating the NIST AI RMF functions as sequential project phases, when GOVERN is cross-cutting and MEASURE and MANAGE activities can occur alongside mapping.

90
MCQhard

A media company serves personalized article recommendations through a model hosted on a cloud inference service. During a major news event, request volume spikes tenfold and p95 latency rises from 120 ms to over 2 seconds, causing timeouts on the web front end. The model itself is unchanged and the endpoint is healthy. The team wants to keep serving personalized results during spikes without degrading the user experience. Which action should the team take first?

A.Retrain the recommendation model on a larger dataset so it produces results faster under load.
B.Switch the front end to a static, non-personalized article list whenever request volume exceeds the autoscaling limit.
C.Lower the model's prediction confidence threshold so fewer candidates are scored per request.
D.Cache or precompute recommendations for high-traffic article and user segments and serve those cached results during the spike.
AnswerD

Serving precomputed or cached recommendations for the hottest segments removes most inference calls from the critical path during the spike, directly reducing p95 latency and preventing front-end timeouts. This is the fastest operational lever because it requires no model change and can be enabled immediately, preserving personalization quality for the segments that matter most during the event.

Why this answer

The latency spike is caused by a tenfold surge in request volume against a fixed-capacity inference path, not by a model defect. Caching or precomputing recommendations for high-traffic segments removes redundant inference work from the request path, cutting p95 latency immediately while preserving personalization. Retraining, threshold tuning, and static fallbacks either do not reduce per-request compute or sacrifice the personalization the team must retain.

Exam trap

The trap here is treating a capacity and serving-path problem as a model-quality problem, which leads to retraining or threshold changes that cannot reduce inference latency.

91
MCQmedium

A company uses an AI model to screen job applications. The model is trained on historical hiring data that reflects past biases. After deployment, the model disproportionately rejects candidates from certain demographics. Which concept does this best illustrate?

A.Overfitting
B.Model drift
C.Algorithmic bias
D.Underfitting
AnswerC

Training on historical hiring data encodes past human decisions, so the model reproduces and amplifies those demographic disparities. This is algorithmic bias: systematic unfair outcomes arising from biased training data rather than explicit discriminatory rules.

Why this answer

Algorithmic bias refers to systematic and unfair discrimination in AI outputs due to biased training data or model design. Option A (overfitting) is about a model that performs well on training data but poorly on new data due to excessive complexity. Option B (model drift) is about performance degradation over time due to changes in data distribution.

Option D (underfitting) is when a model is too simple to capture patterns.

92
MCQmedium

A company wants to automatically group customer support tickets into categories (e.g., billing, technical, account) without pre-labeled data. Which machine learning approach should they use?

A.Supervised classification with logistic regression
B.Semi-supervised learning with a small labeled set
C.Unsupervised clustering using K-means
D.Reinforcement learning with a reward function
AnswerC

K-means partitions unlabelled tickets into k clusters by minimising within-cluster variance across feature vectors, directly satisfying the no-pre-labelled-data constraint. Unlike supervised classification, it discovers the billing, technical and account groupings from inherent similarity rather than predefined labels, matching the stated requirement for automatic categorisation.

Why this answer

The company has no pre-labeled data, which means supervised learning (which requires labeled examples) is not feasible. Unsupervised clustering, such as K-means, groups data points into clusters based on feature similarity without needing any labels, making it ideal for automatically discovering categories like billing, technical, or account from raw ticket text.

Exam trap

CompTIA AI often tests the distinction between supervised and unsupervised learning by presenting a scenario with 'no pre-labeled data' to trick candidates into choosing semi-supervised learning (Option B) because it sounds like a compromise, but the correct answer is always unsupervised clustering when zero labels are available.

How to eliminate wrong answers

Option A is wrong because supervised classification with logistic regression requires a pre-labeled training dataset, which the company does not have. Option B is wrong because semi-supervised learning still requires at least a small set of labeled data to guide the model, contradicting the 'without pre-labeled data' condition. Option D is wrong because reinforcement learning uses a reward function to learn a policy through trial-and-error interactions with an environment, which is not suited for static grouping of text data into categories.

93
MCQeasy

In unsupervised learning, which task involves grouping similar data points together based on feature similarities?

A.Anomaly detection
B.Classification
C.Clustering
D.Regression
AnswerC

Clustering partitions unlabelled data into groups whose members share similar feature values, maximising intra-cluster similarity. It is the unsupervised task defined by grouping similar points, distinguishing it from dimensionality reduction or density estimation, which do not produce discrete similarity-based groupings.

Why this answer

Clustering is the unsupervised learning task that groups similar data points based on feature similarities, without predefined labels. It identifies inherent structures in data by minimizing intra-cluster distances and maximizing inter-cluster distances. Common algorithms include K-means, DBSCAN, and hierarchical clustering.

This contrasts with supervised tasks like classification and regression, which require labeled data.

Exam trap

AI0-001 often tests the distinction between supervised and unsupervised learning tasks, and candidates may confuse clustering with classification because both involve grouping, but classification requires labeled data.

How to eliminate wrong answers

Option A is wrong because anomaly detection identifies rare or unusual data points that deviate from the norm, not grouping similar points together. Option B is wrong because classification is a supervised learning task that assigns predefined labels to data points, requiring labeled training data. Option D is wrong because regression is a supervised learning task that predicts continuous numerical values, not grouping similar data points.

94
MCQhard

A team wants to deploy a large language model on edge devices with limited memory and compute. They need to reduce model size by at least 50% while preserving accuracy. Which combination of techniques is most effective?

A.Apply INT8 quantization and weight pruning
B.Distill the model into a smaller architecture without quantization or pruning
C.Use FP32 precision and increase batch size
D.Use FP16 quantization and add more layers
AnswerA

INT8 quantization shrinks weights from 32-bit to 8-bit, cutting size roughly 75%, while weight pruning removes redundant connections. Combined, they exceed the 50% reduction target with minimal accuracy loss, fitting edge memory and compute limits.

Why this answer

INT8 quantization reduces the precision of weights and activations from 32-bit to 8-bit, cutting memory usage by approximately 75% for those tensors, while weight pruning removes redundant connections, often achieving over 50% size reduction with minimal accuracy loss when combined. Together, they directly address the constraints of edge devices by shrinking the model footprint and computational requirements without requiring a complete architecture redesign.

Exam trap

The exam often tests the misconception that a single technique (like distillation or FP16) is sufficient for aggressive size reduction, when in reality, combining complementary compression methods (quantization and pruning) is necessary to meet both the 50% size reduction and accuracy preservation requirements on edge devices.

How to eliminate wrong answers

Option B is wrong because knowledge distillation alone reduces model size by training a smaller student network, but without quantization or pruning, the student model may still exceed the 50% reduction target or suffer significant accuracy loss if the architecture is not aggressively compressed. Option C is wrong because using FP32 precision and increasing batch size actually increases memory and compute demands, making it unsuitable for resource-constrained edge devices. Option D is wrong because FP16 quantization provides only a 50% memory reduction (not guaranteed to meet the target when combined with adding layers, which increases model size and complexity, often negating the quantization benefit and degrading accuracy on edge hardware without native FP16 support.

95
MCQmedium

A self-driving car uses an AI model that learns by trial and error, receiving rewards for correct actions and penalties for mistakes. This type of learning is:

A.Supervised learning
B.Unsupervised learning
C.Transfer learning
D.Reinforcement learning
AnswerD

Reinforcement learning trains an agent through reward and penalty signals from interacting with its environment, precisely matching the trial-and-error mechanism described. The car's model optimises actions by maximising cumulative reward, satisfying the stem's constraint of learning from consequences rather than labelled examples.

Why this answer

Reinforcement learning (RL) is the correct answer because the self-driving car's AI model learns through trial and error, receiving rewards for correct actions and penalties for mistakes. This feedback-driven process, where an agent interacts with an environment to maximize cumulative reward, is the defining characteristic of reinforcement learning, not supervised or unsupervised learning.

Exam trap

CompTIA often tests the distinction between reinforcement learning and supervised learning by describing a scenario with feedback (rewards/penalties) but no labeled dataset, leading candidates to mistakenly choose supervised learning because they associate 'feedback' with 'labels'.

How to eliminate wrong answers

Option A is wrong because supervised learning requires labeled input-output pairs (e.g., images tagged with 'stop sign') to train a model, not trial-and-error feedback. Option B is wrong because unsupervised learning finds hidden patterns in unlabeled data (e.g., clustering sensor readings) without any reward or penalty signals. Option C is wrong because transfer learning applies knowledge from a pre-trained model to a new but related task, not learning from scratch via rewards and punishments.

96
Multi-Selectmedium

A company is building an AI-based resume screening tool. They want to ensure the system is secure against data poisoning attacks during the training phase. Which THREE of the following are appropriate defensive measures?

Select 3 answers
A.Apply input sanitization to inference-time queries
B.Use robust statistical methods (e.g., trimmed mean) that are less sensitive to outliers
C.Validate and clean training data to remove anomalies and outliers
D.Restrict training data sources to trusted, verified providers only
E.Implement differential privacy during model training
AnswersB, C, D

Trimmed mean aggregation discards extreme values before averaging, directly limiting the influence any single poisoned training sample can exert on model parameters. This satisfies the stem's training-phase constraint by reducing outlier sensitivity, so injected malicious data cannot skew the learned decision boundary.

Why this answer

Option B is correct because robust statistical methods such as trimmed mean, median, or RANSAC reduce the influence of maliciously injected outlier samples during training, directly mitigating data poisoning. Option C is correct because validating and cleaning training data to detect and remove anomalies, label inconsistencies, and outliers prevents poisoned samples from entering the training set in the first place. Option D is correct because restricting training data to trusted, verified providers reduces the attack surface by ensuring provenance and integrity of the data supply chain, which is a key defense against poisoning.

Option A is not appropriate here because input sanitization at inference time addresses runtime adversarial inputs (e.g., evasion or prompt injection), not training-phase poisoning. Option E is not appropriate because differential privacy protects against privacy leakage of individual training records and does not by itself defend against data poisoning attacks.

Exam trap

The AI0-001 exam often tests the distinction between training-phase attacks (data poisoning) and inference-phase attacks (evasion), so candidates mistakenly apply inference-time defenses like input sanitization to training security.

97
MCQmedium

A company using an AI-based hiring tool receives a candidate request for explanation of an automated rejection. Which GDPR principle is most directly relevant?

A.Right to erasure
B.Right to data portability
C.Right to access
D.Right to explanation
AnswerD

Automated rejection decisions fall under GDPR Article 22, which grants data subjects the right to meaningful information about the logic involved and to contest the outcome. The right to explanation directly satisfies the candidate's request for the reasoning behind the hiring tool's decision.

Why this answer

GDPR Article 22 grants individuals the right not to be subject to a decision based solely on automated processing that produces legal or similarly significant effects, and Recital 71 explicitly references the right to obtain an explanation of such decisions. A hiring rejection is a significant decision, so the candidate's request for explanation maps directly to the right to explanation (often framed as part of Article 22 safeguards).

Exam trap

AI0-001 often tests the confusion between the right to access (Article 15) and the right to explanation (Article 22/Recital 71), leading candidates to pick 'right to access' when the question specifically asks about explaining an automated decision.

How to eliminate wrong answers

Option A is wrong because the right to erasure (Article 17) concerns deleting personal data, not obtaining an explanation of an automated decision. Option B is wrong because data portability (Article 20) concerns receiving and transmitting personal data in a machine-readable format, unrelated to decision explanations. Option C is wrong because the right to access (Article 15) lets individuals obtain a copy of their data and certain information, but the specific request for an explanation of automated decision logic is captured by Article 22/Recital 71, not Article 15 alone.

98
MCQhard

A company trains a sentiment analysis model on customer reviews. An attacker submits hundreds of reviews with the word 'excellent' attached to negative feedback, causing the model to classify negative reviews as positive. This is an example of which attack?

A.Data poisoning
B.Model extraction
C.Adversarial example
D.Prompt injection
AnswerA

Data poisoning corrupts the training set, so injecting mislabelled reviews teaches the sentiment model to associate 'excellent' with positive output. This differs from evasion, which manipulates inputs at inference time. The attacker alters learned parameters, satisfying the stem's training-time manipulation constraint.

Why this answer

Data poisoning occurs when an attacker deliberately corrupts the training data to manipulate the model's behavior. By injecting hundreds of reviews that pair the word 'excellent' with negative sentiment, the attacker shifts the model's learned decision boundary, causing it to misclassify genuinely negative reviews as positive. This directly undermines the integrity of the training dataset, which is the hallmark of a data poisoning attack.

Exam trap

The AI0-001 exam often tests the distinction between attacks that occur during training (data poisoning) versus attacks that occur during inference (adversarial examples), so candidates mistakenly choose adversarial example because they focus on the input manipulation rather than the stage of the attack lifecycle.

How to eliminate wrong answers

Option B is wrong because model extraction involves querying a model to reconstruct its parameters or architecture, not corrupting its training data. Option C is wrong because adversarial examples are crafted inputs that fool a trained model at inference time, not during training. Option D is wrong because prompt injection targets large language models by manipulating input prompts to override instructions, not by corrupting training data.

99
Multi-Selectmedium

A company is training a model on proprietary data and wants to prevent data poisoning. Which TWO practices are most important? (Select TWO.)

Select 2 answers
A.Implementing access controls on the training dataset
B.Validating the integrity of training data
C.Using a larger model
D.Increasing training epochs
E.Using homomorphic encryption
AnswersA, B

Access controls restrict who can write to or modify the training dataset, preventing unauthorised actors from injecting malicious samples. This directly addresses the data poisoning threat by limiting the attack surface to trusted contributors, satisfying the stem's requirement to protect proprietary training data.

Why this answer

Option A (Implementing access controls on the training dataset) is correct because data poisoning requires an adversary to inject or modify training samples, and strict authentication/authorization (e.g., IAM roles, least-privilege permissions on the S3 bucket or data lake) prevents unauthorized parties from tampering with the proprietary dataset in the first place. Option B (Validating the integrity of training data) is correct because even with access controls, data can be corrupted or subtly altered, so integrity checks such as cryptographic hashes/checksums, provenance tracking, and outlier or anomaly detection help detect poisoned or tampered samples before they influence the model. Option C (Using a larger model) does not belong because model capacity has no bearing on whether poisoned data enters the pipeline and can even make a model more susceptible to memorizing malicious samples.

Option D (Increasing training epochs) does not belong because more training iterations only reinforce whatever data is present, potentially amplifying the effect of poisoned samples rather than preventing them. Option E (Using homomorphic encryption) does not belong because it protects data confidentiality during computation, not the authenticity or integrity of the training data against poisoning.

Exam trap

The AI0-001 exam often tests the distinction between security controls that prevent attacks (access controls, integrity validation) versus performance tuning (model size, epochs) or privacy techniques (homomorphic encryption), leading candidates to confuse data poisoning prevention with unrelated optimizations.

100
Multi-Selectmedium

Which TWO of the following are appropriate uses of unsupervised learning?

Select 2 answers
A.Classifying emails as spam or not spam
B.Predicting the sale price of a house given its features
C.Detecting unusual patterns in network traffic that may indicate a cyberattack
D.Identifying a person from a photo
E.Segmenting customers into groups based on purchasing behavior
AnswersC, E

Anomaly detection flags deviations from learned normal behaviour without labels, so unlabelled network traffic can be modelled and outliers surfacing as possible cyberattacks. This satisfies unsupervised learning's defining constraint: no target labels are required during training.

Why this answer

Option C is correct because anomaly detection in network traffic is a classic unsupervised learning task: the model is trained on normal traffic without labels and flags deviations that may indicate a cyberattack, using techniques like clustering or autoencoders. Option E is correct because customer segmentation is typically performed with unsupervised clustering algorithms such as k-means or hierarchical clustering, which group customers by purchasing behavior without predefined labels. Option A is incorrect because spam classification is supervised learning, requiring labeled examples of spam and non-spam emails.

Option B is incorrect because predicting a house's sale price is supervised regression with labeled target values. Option D is incorrect because identifying a person from a photo is supervised classification (or a related supervised recognition task) trained on labeled images of known individuals.

Exam trap

CompTIA often tests the distinction between supervised and unsupervised learning by presenting tasks that seem 'automatic' but actually require labeled data, tricking candidates into choosing supervised tasks as unsupervised uses.

101
MCQhard

The exhibit shows the output of a drift monitoring command for a fraud detection model. The team has an automated pipeline that triggers retraining when the overall average drift score exceeds 0.10. Based on the exhibit, what should the operations team do next?

A.Force retraining on all features to ensure the model adapts to the new data distribution.
B.Manually analyze the drift in 'amount' and 'location' and investigate potential causes.
C.No action is needed because the model is performing within acceptable drift limits.
D.Initiate the automated retraining pipeline since the average drift exceeds 0.05.
AnswerB

The exhibit shows per-feature drift exceeding the threshold on 'amount' and 'location' while the overall average stays below 0.10, so the automated retraining trigger never fires. Manual investigation of those two features is required to determine whether the drift is genuine and warrants intervention.

Why this answer

The exhibit shows that the overall average drift score is below 0.10, so the automated retraining pipeline should not trigger. However, individual features like 'amount' and 'location' show elevated drift values that warrant manual investigation to understand root causes before any retraining decision. The team should analyze these specific features to determine if the drift is due to genuine data distribution changes or data quality issues.

Exam trap

CompTIA AI often tests the distinction between aggregate drift thresholds and per-feature drift analysis, trapping candidates who assume that a low overall average drift means no action is needed, ignoring that individual features may still require investigation.

How to eliminate wrong answers

Option A is wrong because force retraining on all features is an overreaction and could introduce model instability; retraining should be targeted or based on the overall drift threshold, not forced indiscriminately. Option C is wrong because while the overall average drift is below 0.10, the elevated drift in 'amount' and 'location' indicates that action is needed to investigate potential causes, so 'no action' is not appropriate. Option D is wrong because the automated retraining pipeline triggers when the overall average drift score exceeds 0.10, not 0.05; the exhibit shows the average drift is below 0.10, so the pipeline should not be initiated.

102
MCQeasy

A data scientist is deploying a machine learning model to production. The model was trained on an imbalanced dataset. Which technique should be used during deployment to mitigate bias without retraining the model?

A.Apply post-processing calibration to adjust decision thresholds
B.Use an ensemble of models trained on balanced subsets
C.Rebalance the dataset using SMOTE before inference
D.Remove sensitive features from the input data
AnswerA

Post-processing calibration adjusts decision thresholds after inference, shifting the operating point to equalise outcomes across groups. Because it modifies predictions rather than learned weights, it mitigates bias from the imbalanced training set without retraining the model.

Why this answer

Post-processing calibration adjusts the decision threshold of the model to account for the class imbalance present in the training data. This technique modifies the output probabilities or classification boundary without requiring access to the original training data or retraining the model, making it suitable for deployment scenarios where the model is already fixed.

Exam trap

CompTIA often tests the distinction between techniques applied during training versus deployment, and the trap here is that candidates mistakenly choose SMOTE or ensemble methods, which require retraining, instead of recognizing that threshold adjustment is a valid post-deployment bias mitigation strategy.

How to eliminate wrong answers

Option B is wrong because using an ensemble of models trained on balanced subsets requires retraining or modifying the model architecture, which violates the constraint of not retraining the model. Option C is wrong because SMOTE (Synthetic Minority Over-sampling Technique) is a data preprocessing method applied before training to balance the dataset, not during inference; applying it at inference time would require access to the original training data and would alter the input distribution, which is not feasible or correct. Option D is wrong because simply removing sensitive features does not mitigate bias caused by imbalanced data; bias can still propagate through correlated features, and this approach does not address the class imbalance issue directly.

103
MCQhard

A media company serves personalized article recommendations through a model that is retrained weekly. After a major news event, engagement metrics show that recommendations became stale within hours because the model had not yet seen the new topic. The engineering team wants recommendations to reflect breaking topics within minutes without retraining the whole model. Which approach should the team implement?

A.Add a real-time retrieval layer that surfaces recently published articles by semantic similarity to the user's current session, and blend those candidates with the model's ranked output.
B.Shorten the retraining cadence from weekly to hourly so the model learns the new topic sooner.
C.Increase the weight of the recency feature in the existing ranking model and redeploy the updated weights.
D.Lower the diversity threshold so the recommender is allowed to show a wider spread of topics to each user.
AnswerA

Real-time retrieval injects brand-new articles into the candidate set immediately, independent of the weekly training cycle, and semantic similarity keeps them relevant to the user's live session. Blending retrieved candidates with the model's ranking preserves personalization quality while closing the freshness gap that pure batch retraining cannot address within minutes.

Why this answer

The gap is that breaking articles never enter the candidate set until the next weekly training cycle, so no amount of reweighting or diversity tuning can surface them. A real-time retrieval layer that matches fresh articles to the live session and blends them into the ranking closes that gap within minutes while preserving existing personalization.

Exam trap

The trap here is trying to solve a candidate-generation freshness problem with ranking-side changes such as feature weights or diversity thresholds, which only reorder items that were already retrieved.

104
MCQmedium

A team is developing a recommendation system for an e-commerce platform. They want to use collaborative filtering but are concerned about cold-start problems for new users. Which approach would best mitigate the cold-start problem?

A.Incorporate user demographic features as side information
B.Increase the number of latent factors in matrix factorization
C.Use a popularity-based baseline for all recommendations
D.Use only item-based collaborative filtering
AnswerA

Demographic side information lets the system infer preferences for users lacking interaction history, so recommendations are generated from attributes rather than behaviour. This directly mitigates the cold-start constraint, where collaborative filtering alone cannot compute similarity for new users.

Why this answer

Incorporating user demographic features as side information allows the collaborative filtering model to generate initial recommendations for new users based on their demographic profile, effectively addressing the cold-start problem. This approach uses content-based features to bootstrap the recommendation process until sufficient user interaction data is collected.

Exam trap

The CompTIA AI+ exam often tests the misconception that increasing model complexity or switching between collaborative filtering variants alone can solve the cold-start problem, when in fact the solution requires incorporating auxiliary data (side information) to bootstrap recommendations for new users.

How to eliminate wrong answers

Option B is wrong because increasing the number of latent factors in matrix factorization does not solve the cold-start problem; it only increases model complexity and may lead to overfitting without addressing the lack of user interaction data. Option C is wrong because using a popularity-based baseline for all recommendations ignores personalization entirely, which defeats the purpose of collaborative filtering and does not leverage user-specific signals even when they become available. Option D is wrong because using only item-based collaborative filtering still requires some user interaction data to compute item similarities; it does not inherently handle new users with no history.

105
MCQhard

A financial institution uses a deep learning model for loan approvals. Under the EU AI Act, this is considered a high-risk AI system. Which mandatory requirement must the institution fulfill before deployment?

A.Obtain certification from an ISO 27001 auditor
B.Publish the model's source code publicly
C.Register the AI system with the national data protection authority
D.Conduct a risk assessment and bias testing
AnswerD

High-risk classification under the EU AI Act obliges providers to establish a risk management system and test for biased outcomes before deployment. This satisfies the stem's pre-deployment constraint, unlike post-market monitoring or voluntary transparency disclosures.

Why this answer

Under the EU AI Act, high-risk AI systems must undergo a conformity assessment that includes a risk assessment and bias testing to ensure fairness, transparency, and non-discrimination before deployment. This requirement is mandated by Articles 9 and 10 of the Act, which specifically address risk management and data governance for high-risk systems. Option D correctly identifies this mandatory step, as the institution must demonstrate that the model does not produce biased outcomes that could lead to discriminatory lending practices.

Exam trap

The AI0-001 exam often tests the misconception that all AI systems require public transparency or external certification, but the EU AI Act specifically mandates internal risk and bias assessments for high-risk systems, not broad publication or ISO standards.

How to eliminate wrong answers

Option A is wrong because ISO 27001 certification pertains to information security management systems, not to AI-specific risk or bias compliance under the EU AI Act; the Act does not require ISO 27001 certification for high-risk AI systems. Option B is wrong because the EU AI Act does not mandate public disclosure of source code; doing so could violate trade secrets and intellectual property rights, and transparency requirements are limited to documentation and logging, not open-source publication. Option C is wrong because registration with a national data protection authority is not a pre-deployment requirement for high-risk AI systems under the EU AI Act; instead, the Act requires registration in an EU-wide database managed by the European Commission, not individual national authorities.

106
MCQmedium

A machine learning engineer is building a model to predict whether a customer will make a purchase within the next week. The dataset contains 10,000 samples with 20 features, and the target variable is binary. The engineer wants to use a model that provides interpretable results to explain predictions to business stakeholders. Which model is most appropriate?

A.Gradient boosting machine
B.Logistic regression
C.Random forest
D.Support vector machine with RBF kernel
AnswerB

Logistic regression is a linear model that provides coefficients for each feature, indicating the direction and magnitude of their influence on the predicted probability. This makes it highly interpretable, as stakeholders can understand how each feature contributes to the prediction. It is well-suited for binary classification tasks and performs reasonably well when the relationship between features and the log-odds is approximately linear. In this scenario, with 20 features and 10,000 samples, logistic regression can be trained efficiently and the coefficients can be explained to business users.

Why this answer

Logistic regression is a linear model that provides coefficients for each feature, making it highly interpretable. Business stakeholders can understand how each feature influences the predicted probability of purchase. While ensemble methods like random forest and gradient boosting may offer higher accuracy, they are less transparent.

Support vector machines with non-linear kernels are also difficult to interpret. Therefore, logistic regression is the most appropriate choice when interpretability is a primary requirement.

Exam trap

The trap here is assuming that more complex models always yield better business value, overlooking the need for interpretability in stakeholder communication.

107
Multi-Selectmedium

A healthcare organization is deploying an AI model to predict patient readmission risk. They must comply with regulations that protect patient privacy. Which TWO techniques should they implement to enhance privacy preservation?

Select 2 answers
A.Data augmentation
B.Differential privacy
C.Model quantization
D.Federated learning
E.Dropout regularization
AnswersB, D

Differential privacy adds calibrated statistical noise to query outputs or training gradients, mathematically bounding how much any single patient's record can influence the model. This directly satisfies the healthcare organisation's regulatory privacy constraint, since attackers cannot reliably infer whether a specific individual's data was included in the readmission-risk training set.

Why this answer

Differential privacy (B) is correct because it adds calibrated statistical noise (e.g., via mechanisms like the Laplace or Gaussian mechanism) to queries or training updates, providing a mathematically provable guarantee that any single patient's data cannot be distinguished in the model's output, which directly supports regulatory privacy requirements such as HIPAA. Federated learning (D) is correct because it trains the model across distributed data sources—such as individual hospitals or devices—keeping patient records local and exchanging only model updates or gradients rather than raw protected health information, thereby minimizing data exposure. The remaining options do not provide privacy guarantees: data augmentation (A) merely expands the training set with synthetic or transformed samples, model quantization (C) reduces numerical precision to shrink model size and speed inference, and dropout regularization (E) randomly deactivates neurons to reduce overfitting—none of these prevent the model from memorizing or leaking sensitive patient information.

Exam trap

The AI0-001 exam often tests the misconception that any regularization or optimization technique (like dropout or quantization) can provide privacy, when in fact only methods that explicitly limit information leakage (like differential privacy and federated learning) are designed for that purpose.

108
Multi-Selecthard

Which THREE are effective methods for ensuring data privacy in AI training? (Choose three.)

Select 3 answers
A.Data encryption at rest
B.Data anonymization
C.Differential privacy
D.Data replication
E.Federated learning
AnswersB, C, E

Data anonymisation removes or alters personally identifiable information before it enters the training set, so the model never learns identifiers tied to real individuals. This directly satisfies the stem's data privacy requirement by preventing re-identification from model outputs or memorised training data, a stronger safeguard than post-training access controls alone.

Why this answer

Data anonymization (B) is correct because removing or masking personally identifiable information (PII) such as names, addresses, and identifiers before training prevents the model from learning or exposing individual identities. Differential privacy (C) is correct because it adds calibrated statistical noise (e.g., via mechanisms like Laplace or Gaussian noise with a privacy budget epsilon) so that the inclusion or exclusion of any single record has a negligible effect on outputs, providing a formal privacy guarantee. Federated learning (E) is correct because it trains models locally on each device or silo and shares only model updates (e.g., gradients or weights) rather than raw data, keeping sensitive data on the originating endpoint.

Data encryption at rest (A) protects stored data against unauthorized access but does not prevent privacy leakage during training or inference, so it is not one of the three methods for ensuring privacy in AI training. Data replication (D) merely copies data to additional locations, which increases exposure and does nothing to protect privacy, so it is not a valid method.

Exam trap

The AI0-001 exam often tests the distinction between security controls (like encryption) and privacy-preserving techniques, trapping candidates who confuse data protection at rest with privacy during model training.

109
MCQmedium

A company is fine-tuning a large language model using PEFT (Parameter-Efficient Fine-Tuning) to reduce GPU memory usage. They have limited hardware and need to fine-tune a 70B parameter model on a single GPU with 24 GB VRAM. Which technique is MOST suitable?

A.Full fine-tuning with gradient checkpointing
B.QLoRA (Quantization-aware LoRA) with 4-bit quantization
C.Instruction tuning with a smaller 7B model
D.LoRA (Low-Rank Adaptation) alone
AnswerB

QLoRA quantises the frozen base weights to 4-bit NF4 and trains only small LoRA adapters, cutting memory enough to fine-tune a 70B model on a single 24 GB GPU. Plain LoRA or full fine-tuning cannot fit within that VRAM budget.

Why this answer

QLoRA combines quantization (4-bit) and LoRA to fine-tune very large models on limited hardware, achieving significant memory reduction while maintaining performance.

110
MCQeasy

A data scientist notices that a binary classification model consistently predicts the majority class. Which data engineering technique should be applied?

A.Feature scaling
B.Dimensionality reduction
C.Polynomial features
D.Oversampling
AnswerD

Consistent majority-class prediction indicates severe class imbalance, where the loss is minimised by always guessing the dominant label. Oversampling replicates or synthesises minority-class examples, rebalancing the training distribution so the classifier learns the minority decision boundary.

Why this answer

Oversampling (Option D) is correct because the model's bias toward the majority class indicates a class imbalance problem. By synthetically increasing the number of minority class samples (e.g., using SMOTE or random oversampling), the training data becomes more balanced, allowing the classifier to learn decision boundaries that are not skewed toward the majority class.

Exam trap

CompTIA often tests the misconception that feature scaling or dimensionality reduction can fix class imbalance, when in reality these techniques address different issues like feature magnitude or curse of dimensionality, not skewed target distributions.

How to eliminate wrong answers

Option A is wrong because feature scaling normalizes the range of input features (e.g., via min-max scaling or standardization) but does not address class imbalance; it only prevents features with larger magnitudes from dominating gradient-based optimization. Option B is wrong because dimensionality reduction (e.g., PCA or t-SNE) reduces the number of features to combat overfitting or noise, but it does not alter the class distribution, so the majority class bias remains. Option C is wrong because polynomial features create interaction or higher-degree terms from existing features to capture non-linear relationships, but they do not change the ratio of majority to minority samples, leaving the imbalance untouched.

111
Multi-Selecthard

An AI engineer is fine-tuning a transformer-based language model for a domain-specific task. They want to improve the model's factual accuracy and reduce hallucinations. Which THREE strategies should they consider? (Select THREE)

Select 3 answers
A.Increase the model's context window size beyond the training limit
B.Fine-tune the model on a curated domain-specific corpus
C.Use a higher temperature setting during generation
D.Apply chain-of-thought prompting for complex queries
E.Implement Retrieval-Augmented Generation (RAG)
AnswersB, D, E

Fine-tuning on a curated domain-specific corpus updates the model's weights toward accurate, in-domain factual patterns, directly reducing hallucination on that task. This satisfies the stem's goal of improving factual accuracy by grounding the model in verified domain content rather than relying on broad pretraining data.

Why this answer

Option B is correct because fine-tuning on a curated, high-quality domain-specific corpus injects accurate, task-relevant knowledge into the model's weights, which directly improves factual grounding and reduces the likelihood of fabricating domain facts. Option D is correct because chain-of-thought prompting encourages the model to decompose complex queries into intermediate reasoning steps, which empirically reduces errors and unsupported claims on multi-step or reasoning-heavy tasks. Option E is correct because Retrieval-Augmented Generation grounds generation in externally retrieved, up-to-date documents at inference time, so the model conditions its output on verifiable evidence rather than relying solely on parametric memory, which is the standard technique for reducing hallucinations.

Option A is not appropriate because simply increasing the context window beyond the training limit is not a supported or effective strategy—models cannot attend beyond their trained positional range without architectural changes or extrapolation techniques, and a larger window alone does not improve factual accuracy. Option C is incorrect because raising the temperature increases sampling randomness and diversity, which typically increases hallucination rather than reducing it; lower temperature is preferred for factual accuracy.

Exam trap

The CompTIA AI+ exam often tests the misconception that increasing randomness (higher temperature) or extending context windows beyond training limits can improve factual accuracy, when in fact these techniques degrade reliability.

112
MCQeasy

A marketing team wants to use a third-party generative AI service to create ad copy. The service provider states that submitted prompts and outputs may be used to improve its models. The company's legal team is concerned about confidential product launch details being entered into the tool. Which of the following is the MOST appropriate first step?

A.Prohibit entering confidential information into the tool until a data processing agreement and enterprise configuration that excludes data from training are in place.
B.Allow the team to proceed because the provider's public privacy policy already covers all customer data handling.
C.Instruct the team to paraphrase confidential details before submitting prompts so the information is no longer recognizable.
D.Require the team to delete the chat history after each session to remove the provider's copy of the data.
AnswerA

The provider's default terms allow submitted content to be used for model improvement, which conflicts with protecting confidential launch details. The immediate control is to stop sensitive input while negotiating contractual protections such as a data processing agreement and enabling an enterprise setting that disables training on submitted data. This addresses the risk at its source before any ad copy is generated.

Why this answer

When a provider's terms permit using inputs for model improvement, confidential material should not be entered until contractual and technical safeguards exist. A data processing agreement plus an enterprise configuration that excludes submissions from training reduces the exposure directly. Privacy policies, paraphrasing, and chat deletion do not create enforceable confidentiality protections, so they fail to address the identified risk.

Exam trap

The trap here is assuming that deleting chat history or paraphrasing prompts removes the provider's ability to retain or learn from submitted content.

113
MCQhard

A team is training a recurrent neural network (RNN) with LSTM units to predict stock prices. The validation loss is significantly higher than the training loss. Which action is MOST likely to reduce the gap?

A.Increase the number of LSTM units
B.Increase the number of training epochs
C.Reduce the sequence length
D.Increase the dropout rate in LSTM layers
AnswerD

Dropout randomly deactivates LSTM units during training, preventing the network from relying on particular hidden-state pathways and reducing variance. The validation loss exceeding training loss indicates overfitting, so increasing dropout regularises the recurrent layers and narrows that generalisation gap.

Why this answer

A significant gap between training and validation loss indicates overfitting, where the model memorizes training data but fails to generalize. Increasing dropout in LSTM layers is a regularization technique that randomly deactivates neurons during training, forcing the network to learn more robust features and reducing overfitting. This directly addresses the gap by improving validation performance.

Exam trap

AI0-001 often tests the diagnosis of overfitting versus underfitting; candidates may pick increasing epochs or units thinking more training will help, but those actions worsen overfitting when the validation loss is already higher than training loss.

How to eliminate wrong answers

Option A is wrong because increasing the number of LSTM units adds model capacity, which would likely worsen overfitting and increase the gap further. Option B is wrong because training for more epochs allows the model to memorize the training data even more, exacerbating overfitting. Option C is wrong because reducing sequence length may lose temporal dependencies and does not directly address overfitting; it could even hurt performance if the shortened sequences lack predictive information.

114
MCQhard

A team is fine-tuning a large language model using LoRA. They have limited GPU memory. Which technique can further reduce memory consumption while maintaining similar fine-tuning quality?

A.Fine-tune all layers instead of using LoRA
B.Increase the rank of LoRA adapters
C.Use QLoRA with 4-bit quantization of the base model
D.Use a larger batch size
AnswerC

QLoRA quantises the frozen base weights to 4-bit NF4 while training LoRA adapters in higher precision, cutting GPU memory far below standard LoRA. This satisfies the limited-memory constraint, and because only adapters are trained, fine-tuning quality stays comparable.

Why this answer

QLoRA extends LoRA by quantizing the frozen base model to 4-bit precision (using NF4 quantization) while keeping LoRA adapters in higher precision, dramatically reducing GPU memory for the base weights. This allows fine-tuning of large models on a single consumer GPU with quality comparable to 16-bit LoRA. The 4-bit base model is dequantized on the fly during forward and backward passes, so memory savings come primarily from storing the base weights in 4 bits instead of 16.

Exam trap

The trap is confusing LoRA with QLoRA — candidates may think LoRA alone already minimizes memory, but LoRA still stores the base model in 16-bit; only QLoRA's 4-bit quantization of the base model provides the additional memory reduction.

How to eliminate wrong answers

Option A is wrong because fine-tuning all layers requires storing full gradients and optimizer states for every parameter, which increases memory usage by orders of magnitude compared to LoRA. Option B is wrong because increasing the LoRA rank adds more trainable parameters and optimizer state, increasing memory consumption, not reducing it. Option D is wrong because a larger batch size increases activation memory and gradient memory, making GPU memory pressure worse, not better.

115
MCQeasy

A data science team is preparing a dataset for a supervised learning task. They split the data into training and test sets. The team then normalizes the features using the mean and standard deviation calculated from the entire dataset before splitting. What issue does this introduce?

A.It improves model generalization
B.It introduces train/test leakage
C.It causes the model to overfit the training data
D.It reduces the variance of the features
AnswerB

Computing the mean and standard deviation across the whole dataset lets statistics from the test set influence the scaling applied to training features. Test information therefore leaks into training, producing optimistic evaluation results that will not generalise to unseen data.

Why this answer

Computing normalization statistics (mean and standard deviation) on the entire dataset before splitting means the test set's statistics influence the training data transformation. This is a form of train/test leakage: information from the test set leaks into the training pipeline, producing an optimistically biased estimate of model performance. The correct approach is to fit the scaler on the training set only and apply the same transformation to the test set.

Exam trap

AI0-001 often tests whether candidates recognize that preprocessing steps (scaling, imputation, encoding) must be fit only on training data — many candidates focus on model training and forget that leakage can occur in the data preparation pipeline.

How to eliminate wrong answers

Option A is wrong because leakage does not improve true generalization — it only inflates reported metrics, which can mislead model selection and deployment decisions. Option C is wrong because overfitting refers to the model memorizing training data; leakage is a data preparation flaw independent of model complexity, and it can occur even with simple models. Option D is wrong because normalization does reduce variance in a mathematical sense, but that is a side effect, not the issue introduced by computing statistics on the full dataset — the actual problem is leakage.

116
MCQmedium

A data science team is deploying a deep learning model for real-time inference on edge devices with limited power and memory. Which model optimisation technique would be MOST effective for reducing latency and memory footprint while maintaining acceptable accuracy?

A.Use a larger batch size during inference
B.Train the model for more epochs to improve convergence
C.Apply quantisation to convert weights from FP32 to INT8
D.Increase the number of layers to improve feature extraction
AnswerC

Quantisation converts FP32 weights to INT8, cutting memory footprint roughly fourfold and enabling integer arithmetic that executes faster on edge hardware, directly satisfying the limited power and memory constraint. Accuracy loss stays acceptable because scaling factors preserve the weight distribution, making it ideal for real-time inference on constrained devices.

Why this answer

Quantization reduces the precision of model weights from 32-bit floating point (FP32) to 8-bit integer (INT8), which directly cuts memory usage by 75% and accelerates inference on edge devices by leveraging integer arithmetic. This technique is specifically designed for resource-constrained environments where power and memory are limited, and it typically preserves accuracy within 1-2% of the original model.

Exam trap

CompTIA often tests the misconception that increasing model complexity (more layers or epochs) improves deployment performance, when in fact the opposite is true for edge inference; candidates may confuse training optimization with inference optimization.

How to eliminate wrong answers

Option A is wrong because increasing batch size during inference increases memory consumption and latency on edge devices, as it requires processing multiple inputs simultaneously, which is counterproductive for real-time, low-latency requirements. Option B is wrong because training for more epochs improves convergence and accuracy but does not reduce model size or inference latency; it may even lead to overfitting without any benefit to deployment efficiency. Option D is wrong because adding more layers increases the model's parameter count, memory footprint, and computational latency, directly opposing the goal of reducing resource usage on edge devices.

117
MCQhard

A financial services firm deploys a credit-scoring model that must produce explanations for adverse action notices. The compliance team requires that each decision be traceable to the exact model version, input features, and the explanation method used at inference time. The data science team currently logs only predictions and timestamps. Which approach best satisfies the traceability requirement?

A.Enable verbose application logging that records request IDs and HTTP status codes for all inference calls.
B.Log the prediction, the model version identifier, a hash of the input feature vector, and the explanation output for every inference request in an immutable audit store.
C.Retrain the model monthly and store the training dataset alongside the production model artifact.
D.Use a model registry to promote models to production and tag each release with a semantic version.
AnswerB

This captures the model version, the exact inputs, and the explanation produced at decision time, which is what the compliance team requires. Storing them immutably supports audits and reproducibility. A feature-vector hash allows verification without retaining raw sensitive data, and the explanation output ties the notice to the actual method used.

Why this answer

Traceability for adverse action notices requires linking each decision to the model version, the exact inputs, and the explanation method. Logging those elements together in an immutable store creates an auditable record. Registry tags, training data retention, and generic request logs each capture only part of the picture and cannot reconstruct an individual decision.

Exam trap

The trap here is confusing model governance artifacts, such as a registry or training data, with per-decision audit records that explain a specific outcome.

118
MCQmedium

A data scientist needs to deploy a PyTorch model to production with low-latency inference. The model must be served as a REST API and should support GPU acceleration. Which combination of tools is MOST suitable for this task?

A.ONNX runtime with a gRPC endpoint on a CPU-only node
B.Apache Spark with MLlib to serve the model in batch mode
C.Docker container with a FastAPI application and Nvidia GPU support
D.Kubeflow Pipelines to deploy the model as a scheduled job
AnswerC

Docker packages the model and FastAPI exposes it as a REST API, while Nvidia GPU support enables CUDA acceleration for low-latency inference. This combination satisfies both the REST API and GPU acceleration constraints for the PyTorch model.

Why this answer

It combines Docker containerization with FastAPI for a lightweight REST API and NVIDIA GPU support (via nvidia-docker or NVIDIA Container Toolkit) to enable low-latency GPU-accelerated inference. This stack directly meets the requirements of low-latency inference, REST API serving, and GPU acceleration without unnecessary overhead.

Exam trap

CompTIA often tests the distinction between batch/offline processing tools (like Spark or Kubeflow Pipelines) and real-time serving frameworks, leading candidates to confuse orchestration or batch tools with low-latency inference solutions.

How to eliminate wrong answers

Option A is wrong because ONNX Runtime with a gRPC endpoint on a CPU-only node cannot provide GPU acceleration, which is explicitly required. Option B is wrong because Apache Spark with MLlib is designed for distributed batch processing and large-scale data pipelines, not for low-latency real-time REST API serving of a single PyTorch model. Option D is wrong because Kubeflow Pipelines is a workflow orchestration tool for scheduling and managing ML pipelines, not a real-time inference serving solution; it lacks native REST API endpoints for low-latency inference.

119
MCQhard

A team trains a recurrent neural network to translate sentences averaging 60 words. During evaluation they notice that translations of the final words in long sentences are frequently wrong, while the opening words are translated accurately. Which architectural change best addresses this behavior?

A.Lower the learning rate and train for more epochs
B.Replace the recurrent layers with an attention-based transformer encoder-decoder
C.Increase the batch size used during training
D.Increase the number of recurrent layers stacked in the encoder
AnswerB

The described degradation at the end of long sequences is characteristic of vanishing gradients and limited memory in plain recurrent networks. Self-attention lets every output position attend directly to every input position in constant path length, so information from early and late tokens remains accessible. Transformers are the standard architecture for long-sequence translation and directly resolve this failure.

Why this answer

Recurrent networks compress the entire source sentence into a fixed-size hidden state and suffer vanishing gradients across many timesteps, which disproportionately harms recall of distant tokens. An attention-based transformer removes the sequential bottleneck, allowing direct token-to-token connections and preserving context for the final words of long sentences.

Exam trap

The trap here is assuming more layers, larger batches, or longer training can overcome a positional memory problem, when the limitation is architectural rather than an optimization or capacity issue.

120
MCQhard

A data scientist trains a deep neural network for image classification. The training loss decreases but validation loss starts increasing after 50 epochs. What should the data scientist do to improve generalization?

A.Decrease batch size
B.Apply dropout and early stopping
C.Add more hidden layers
D.Increase learning rate
AnswerB

Dropout randomly deactivates neurons during training, reducing co-adaptation that drives overfitting, while early stopping halts training once validation loss begins rising. Together they directly counter the divergence at epoch 50, restoring generalisation without altering the model architecture or dataset.

Why this answer

The increasing validation loss while training loss decreases is a classic sign of overfitting. Dropout randomly deactivates neurons during training, which prevents co-adaptation and forces the network to learn more robust features. Early stopping halts training when validation performance stops improving, directly addressing the overfitting by selecting the model with the best generalization before it degrades.

Exam trap

CompTIA often tests the misconception that increasing model complexity (more layers) or adjusting batch size/learning rate can fix overfitting, when in reality these changes either exacerbate the problem or address unrelated training dynamics.

How to eliminate wrong answers

Option A is wrong because decreasing batch size introduces more noise into the gradient estimates, which can actually hurt generalization and may lead to slower convergence or instability, not a direct cure for overfitting. Option C is wrong because adding more hidden layers increases model capacity and complexity, which typically worsens overfitting by allowing the network to memorize the training data even more. Option D is wrong because increasing the learning rate can cause the optimizer to overshoot minima, leading to divergence or poor convergence, and does not address the fundamental issue of the model fitting noise in the training data.

121
MCQmedium

An LLM-based chatbot is being deployed for customer support. The security team wants to prevent the bot from generating toxic or harmful responses. Which defense is MOST appropriate?

A.Input validation and sanitization
B.Rate limiting on API requests
C.Output filtering and guardrails
D.Red teaming the AI system
AnswerC

Output filtering inspects the model's generated text before it reaches the user, blocking toxic or harmful content regardless of how the prompt was phrased. Guardrails enforce policy at that boundary, satisfying the requirement to prevent harmful responses rather than merely discouraging them through input sanitisation.

Why this answer

Output filtering and guardrails can block harmful content before it reaches the user. Input validation sanitizes inputs, red teaming identifies vulnerabilities, and rate limiting prevents abuse but not toxic content.

122
Multi-Selecteasy

An organization is planning to fine-tune an open-source LLM for internal use. To secure the supply chain, which TWO steps should they take before using the base model? (Select two.)

Select 2 answers
A.Retrain the model from scratch
B.Verify the model's provenance and checksums
C.Vet the pre-trained model for potential backdoors
D.Set up audit logging of all interactions
E.Fine-tune the model on sensitive internal data
AnswersB, C

Verifying provenance and checksums confirms the downloaded base model genuinely originates from the trusted publisher and has not been altered in transit or tampered with in the repository, directly addressing supply chain integrity before fine-tuning begins.

Why this answer

Option B is correct because verifying the model's provenance and checksums confirms that the base model was obtained from a trusted source and has not been tampered with or substituted during download, which is a foundational supply-chain control. Option C is correct because pre-trained models can contain hidden backdoors or malicious behaviors (e.g., triggered outputs or poisoned weights), so vetting the model before fine-tuning helps detect such threats prior to integrating it into internal systems. Option A is not appropriate because retraining from scratch is prohibitively expensive and unnecessary for supply-chain security.

Option D is a runtime monitoring control that occurs after deployment, not a pre-use supply-chain step. Option E is incorrect because fine-tuning on sensitive internal data increases risk and does not secure the base model's supply chain.

Exam trap

CompTIA often tests the distinction between pre-deployment supply chain security (verification and vetting) and post-deployment operational controls (logging, fine-tuning), tricking candidates into selecting runtime measures for a supply chain question.

123
MCQhard

A media company uses a natural language processing (NLP) model to classify news articles into topics. The model was trained on articles from 2015-2018. In 2023, the model's F1 score drops significantly. The data scientists find that the word embeddings no longer capture the meaning of some terms (e.g., 'covid', 'metaverse'). The model uses static word embeddings (Word2Vec) trained on the original corpus. Which solution BEST addresses the observed degradation? A. Replace static embeddings with contextual embeddings from a transformer model like BERT, then fine-tune the classifier. B. Retrain the static Word2Vec embeddings on a larger corpus from 2023. C. Apply data augmentation to the original training data by replacing words with synonyms. D. Increase the dimensionality of the static embeddings.

A.Retrain the static Word2Vec embeddings on a larger corpus from 2023.
B.Increase the dimensionality of the static embeddings.
C.Replace static embeddings with contextual embeddings from a transformer model like BERT, then fine-tune the classifier.
D.Apply data augmentation to the original training data by replacing words with synonyms.
AnswerC

Contextual embeddings dynamically represent words based on context, handling semantic shift effectively.

Why this answer

Contextual embeddings (e.g., BERT) capture meaning based on context, adapting to new uses of words like 'covid' meaning pandemic. Fine-tuning the classifier on new data would update the model. Option A (retraining static embeddings) might capture new word senses but still assigns a single vector per word, missing context.

Option B (increasing dimensionality) does not address the semantic shift. Option D (data augmentation) does not introduce new word meanings.

124
Multi-Selectmedium

A data scientist suspects a model extraction attack on their deployed classifier. Which TWO indicators are MOST consistent with such an attack? (Select two.)

Select 2 answers
A.Queries that include SQL injection attempts
B.Queries that repeatedly ask for the same prediction
C.A large number of queries from a single IP address over a short period
D.Queries with special characters attempting to reveal system prompts
E.Queries covering a wide range of diverse inputs
AnswersC, E

Model extraction relies on high-volume automated querying to approximate the decision boundary. A single IP issuing many queries in a short window satisfies the volume and rate constraint that distinguishes extraction from ordinary sporadic user traffic.

Why this answer

Option C is correct because model extraction (model stealing) attacks require the adversary to submit a very high volume of prediction requests to the deployed classifier, often from a single source IP, in order to gather enough input-output pairs to train a substitute model; a burst of queries from one IP in a short window is a classic volumetric indicator. Option E is correct because effective extraction depends on sampling the model's decision space broadly, so queries that systematically span a wide, diverse range of inputs (rather than a narrow slice) are consistent with an attacker trying to approximate the full decision boundary. Option A is not correct because SQL injection targets database-backed application inputs and is unrelated to stealing a model's parameters or behavior.

Option B is not correct because repeatedly requesting the same prediction yields redundant, low-information samples and is more indicative of caching, retry, or probing behavior than systematic extraction. Option D is not correct because attempts to reveal system prompts are prompt-injection or prompt-leaking attacks against LLM applications, not model extraction of a classifier.

Exam trap

CompTIA AI exams often test the distinction between model extraction (which requires diverse inputs to map the decision boundary) and denial-of-service or brute-force attacks (which involve repeated identical queries), so candidates mistakenly select Option B thinking any high query volume indicates extraction.

125
Multi-Selectmedium

A data scientist is evaluating a binary classifier for a medical diagnosis task. The dataset is imbalanced with 5% positive cases. Which THREE metrics should the data scientist consider for a comprehensive evaluation?

Select 3 answers
A.Precision
B.F1 score
C.Accuracy
D.Perplexity
E.Recall
AnswersA, B, E

Precision measures the proportion of predicted positives that are truly positive, exposing how many false alarms the classifier raises. At a 5% positive rate, this matters because accuracy alone would look high while missing poor positive-class performance.

Why this answer

Precision (A) is correct because it measures the proportion of predicted positives that are truly positive, which is critical in an imbalanced medical diagnosis task where false positives carry real cost. Recall (E) is correct because it measures the proportion of actual positives that are correctly identified, ensuring the classifier does not miss the rare 5% positive cases. F1 score (B) is correct because it is the harmonic mean of precision and recall, providing a single balanced metric that is far more informative than accuracy on skewed class distributions.

Accuracy (C) is not appropriate here because a trivial model predicting all negatives would achieve 95% accuracy while detecting zero positive cases. Perplexity (D) does not belong because it is a language-model metric measuring how well a probability distribution predicts a sample, not a classification performance measure.

Exam trap

The trap here is that candidates often default to accuracy as a universal metric, but CompTIA AI tests the understanding that accuracy is unreliable for imbalanced datasets, and that metrics like precision, recall, and F1 score are required for a comprehensive evaluation.

126
Multi-Selecthard

Which TWO of the following are techniques used for reducing overfitting in neural networks? (Choose two.)

Select 2 answers
A.Dropout
B.Boosting
C.L2 regularization
D.Increasing the learning rate
E.Increasing the number of hidden layers
AnswersA, C

Dropout randomly deactivates units each training pass, preventing neurons from co-adapting to noise and forcing redundant representations. This regularising effect reduces overfitting, satisfying the question's requirement for a technique that improves generalisation rather than training accuracy.

Why this answer

Dropout (A) is correct because it randomly deactivates a fraction of neurons during each training iteration, which prevents the network from relying on specific neurons and forces it to learn more robust, generalizable features, thereby reducing overfitting. L2 regularization (C) is correct because it adds a penalty term proportional to the squared magnitude of the weights to the loss function, discouraging large weights and constraining model complexity to improve generalization. Boosting (B) is an ensemble meta-algorithm that combines weak learners to reduce bias, not a neural-network overfitting-reduction technique.

Increasing the learning rate (D) typically causes unstable training or divergence rather than reducing overfitting. Increasing the number of hidden layers (E) raises model capacity and usually worsens overfitting rather than mitigating it.

Exam trap

CompTIA often tests the distinction between regularization techniques and other training strategies, so the trap here is that candidates may confuse boosting (an ensemble method) with regularization, or assume that increasing model complexity (more layers) or learning rate can help reduce overfitting when they actually do the opposite.

127
MCQhard

An organization is developing an AI system to approve loan applications. They want to ensure the model does not discriminate based on race or gender. Which technique BEST addresses this concern?

A.Remove race and gender features from the training data.
B.Use a more complex model to capture nuances.
C.Apply adversarial debiasing during model training.
D.Collect more training data from diverse populations.
AnswerC

Adversarial debiasing trains a predictor alongside an adversary that tries to infer the protected attribute from predictions, forcing representations that cannot distinguish race or gender. This directly reduces disparate impact during training, unlike post-hoc inspection alone.

Why this answer

Adversarial debiasing is a technique that explicitly trains the model to remove sensitive information (like race or gender) from its internal representations, preventing the model from learning discriminatory patterns even if correlated features remain. This directly addresses fairness by making the model's predictions independent of protected attributes, which is more robust than simply removing features (which can still allow proxy discrimination).

Exam trap

CompTIA often tests the misconception that removing protected attributes is sufficient to eliminate bias, when in reality proxy features and correlated variables can still cause discrimination, making adversarial debiasing or other fairness-aware algorithms necessary.

How to eliminate wrong answers

Option A is wrong because simply removing race and gender features does not prevent the model from learning proxies for these attributes (e.g., zip code, income bracket) that can still lead to discriminatory outcomes. Option B is wrong because using a more complex model increases the risk of overfitting to spurious correlations and does not inherently address fairness; it may even amplify biases present in the data. Option D is wrong because collecting more diverse data does not guarantee fairness; biased labeling, historical discrimination, or imbalanced representation can persist, and the model may still learn to discriminate unless debiasing techniques are applied.

128
MCQeasy

A bank uses an AI model to approve loans. During an audit, it is found that the model denies loans at a higher rate for a certain ethnic group. Which governance principle is primarily violated?

A.Accountability
B.Fairness
C.Transparency
D.Privacy
AnswerB

Fairness is violated because the model produces disparate denial rates across ethnic groups, breaching equitable treatment. This directly addresses the stem's constraint: audit-detected bias in loan approvals. Fairness in AI governance requires identifying and mitigating such discriminatory outcomes, ensuring decisions are not systematically disadvantageous to protected groups.

Why this answer

The model's disparate impact on a specific ethnic group directly violates the principle of Fairness, which requires that AI systems do not discriminate based on protected attributes such as race, ethnicity, or gender. In lending, fairness is often assessed using metrics like demographic parity or equal opportunity, and a higher denial rate for one group indicates a lack of algorithmic fairness.

Exam trap

The AI0-001 exam often tests the distinction between Fairness and Transparency, where candidates mistakenly choose Transparency because they think 'explaining the bias' is the primary issue, but the question asks which principle is violated by the biased outcome itself.

How to eliminate wrong answers

Option A is wrong because Accountability refers to the assignment of responsibility for the model's decisions and outcomes, not the presence of bias itself; while the bank must be accountable for the bias, the primary violation here is the discriminatory outcome. Option C is wrong because Transparency concerns the ability to explain and understand how the model makes decisions (e.g., through interpretability or documentation), but the core issue is the biased result, not a lack of explanation. Option D is wrong because Privacy involves the protection of personal data and compliance with regulations like GDPR or CCPA; the scenario does not describe unauthorized data use or exposure, only discriminatory lending decisions.

129
MCQhard

A company uses Azure OpenAI to generate customer support responses. The team notices that repeated queries with similar context incur high costs due to token usage. They want to reduce costs without affecting response quality. Which strategy is MOST effective?

A.Use a larger model to improve efficiency
B.Increase the frequency penalty
C.Reduce the max_tokens parameter
D.Implement prompt caching
AnswerD

Prompt caching stores and reuses the computed key-value representations of repeated context, so identical or similar prefixes are billed at a reduced cached-token rate rather than full input tokens. This directly cuts token costs on recurring queries while returning the same model output, preserving response quality.

Why this answer

Prompt caching stores and reuses tokens from previous queries, reducing token consumption for similar requests and lowering costs without quality loss.

130
MCQeasy

An AI practitioner needs to measure the performance of a binary classification model for disease detection, where the cost of false negatives is very high. Which metric should be prioritized?

A.Recall
B.Precision
C.F1-score
D.Accuracy
AnswerA

Recall measures the proportion of actual positives correctly identified, directly minimising false negatives. In disease detection, missing a genuine case is costly, so maximising recall ensures fewer missed diagnoses, even at the expense of more false positives.

Why this answer

Recall (true positive rate) minimises false negatives, which is critical when missing a positive case is dangerous.

131
MCQmedium

A data pipeline processes customer data from multiple sources. The data quality check reveals duplicate records. Which step should the pipeline include to handle this?

A.Data deduplication
B.Data encryption
C.Data transformation
D.Data validation
AnswerA

Data deduplication removes duplicate records by identifying matching rows across sources and retaining a single canonical entry, directly resolving the duplicate records the quality check flagged. It satisfies the stem's constraint of handling duplicates within the pipeline, unlike profiling or validation, which detect issues but do not eliminate them.

Why this answer

Duplicate records in a data pipeline compromise data integrity and downstream analytics. Data deduplication (Option A) is the correct step because it identifies and removes redundant entries based on key fields or fuzzy matching, ensuring each customer record is unique. This is a core data quality operation in ETL pipelines, often implemented via hash-based comparison or SQL window functions like ROW_NUMBER().

Exam trap

This question tests the distinction between data quality actions (deduplication) and data security or formatting actions (encryption, transformation), leading candidates to confuse validation (which only flags issues) with remediation (which removes duplicates).

How to eliminate wrong answers

Option B (Data encryption) is wrong because encryption secures data at rest or in transit (e.g., AES-256, TLS 1.3) and does not address duplicate records. Option C (Data transformation) is wrong because transformation changes data format or structure (e.g., type casting, normalization) but does not inherently remove duplicates. Option D (Data validation) is wrong because validation checks data against rules (e.g., schema constraints, range checks) and flags errors, but it does not actively eliminate duplicate rows.

132
MCQmedium

A data science team is training an image classification model for a medical imaging application. To prevent data leakage, they must partition the dataset correctly. Which approach ensures that no patient images appear in both training and test sets?

A.Split by patient ID so that all images of a patient go to one set only
B.Shuffle the dataset and take the first 80% for training and last 20% for testing
C.Use k-fold cross-validation without grouping
D.Randomly split all images into training and test sets
AnswerA

Grouping all images from one patient into a single partition prevents the same patient's anatomy appearing in both training and test sets. Patient ID is the grouping key that satisfies the stem's no-overlap constraint, since random image-level splitting would leak patient-specific features.

Why this answer

Data leakage occurs when information from the test set leaks into training. Splitting by patient ID ensures that all images from the same patient are kept together in one partition.

133
MCQeasy

An AI developer needs to store large amounts of unstructured data (e.g., images, logs) for training datasets. Which cloud storage solution is purpose-built for data lakes?

A.Amazon DynamoDB
B.Amazon RDS
C.Amazon S3
D.Amazon Redshift
AnswerC

Amazon S3 provides object storage with flat namespace, unlimited scalability and high durability, purpose-built for data lakes holding unstructured images and logs. It satisfies the stem's requirement for storing large volumes of unstructured training data, unlike block or file storage. S3's decoupling of metadata from objects suits analytics engines querying datasets directly.

Why this answer

Amazon S3 is purpose-built for data lakes because it provides virtually unlimited scalability, high durability (99.999999999% or 11 nines), and supports any type of unstructured data (images, logs, videos) with a flat object storage architecture. Its integration with AWS Glue, Athena, and Lake Formation enables schema-on-read analytics, making it the foundational service for building a data lake on AWS.

Exam trap

Candidates often mistake Redshift for a data lake solution because they associate 'data' with 'warehouse' rather than recognizing that data lakes require raw object storage.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for low-latency transactional workloads, not for storing large volumes of unstructured data for analytics. Option B is wrong because Amazon RDS is a relational database service for structured data with fixed schemas, and it cannot scale to petabyte-scale unstructured data storage. Option D is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for structured, columnar data and SQL-based analytics, not for storing raw unstructured data like images or logs.

134
Multi-Selecteasy

A data engineer is preparing a dataset for a binary classification model. The dataset has 10,000 samples with 100 features. To improve model performance and reduce training time, the engineer decides to perform feature selection. Which two techniques are appropriate for this task? (Select TWO).

Select 2 answers
A.Normalization
B.Recursive Feature Elimination (RFE)
C.L1 Regularization
D.One-Hot Encoding
E.Principal Component Analysis (PCA)
AnswersB, C

Recursive Feature Elimination fits a model, ranks features by importance, removes the weakest, and repeats, progressively shrinking the 100-feature set. This reduces training time and can improve performance by eliminating irrelevant or noisy predictors from the dataset.

Why this answer

Recursive Feature Elimination (RFE) (B) is a wrapper-based feature selection method that repeatedly trains a model, ranks features by importance (e.g., coefficients or feature importances), and prunes the least important ones until the desired number of features remains, directly reducing the 100-feature space to improve performance and cut training time. L1 Regularization (C) — Lasso — adds a penalty equal to the absolute value of coefficients to the loss function, which drives many feature coefficients exactly to zero, effectively performing embedded feature selection and yielding a sparse model. Normalization (A) is a scaling preprocessing step (e.g., min-max or z-score) that changes feature magnitudes but does not remove features, so it is not feature selection.

One-Hot Encoding (D) is a categorical-encoding transformation that expands categorical variables into binary columns, increasing dimensionality rather than reducing it. Principal Component Analysis (E) is a dimensionality-reduction technique that creates new uncorrelated components from linear combinations of the original features, but it is not feature selection because it does not retain or select the original features.

Exam trap

CompTIA often tests the distinction between feature selection (keeping original features) and dimensionality reduction (creating new features), so candidates mistakenly select PCA thinking it selects features, when it actually transforms them into principal components.

135
MCQmedium

Based on the exhibit, what is the most likely cause of the pod failure and its solution?

A.The node has insufficient CPU; add more CPU.
B.The pod is configured with wrong GPU drivers; update drivers.
C.The model is too large; use a smaller model.
D.The container memory limit is too low; increase the memory limit in the pod spec.
AnswerD

The container exceeded its configured memory limit, so the kernel terminated it with an OOMKilled status. Raising the memory limit in the pod spec allows the workload to complete, satisfying the resource constraint the exhibit shows the container breaching.

Why this answer

The pod failure is caused by an OOMKilled (Out of Memory) error, as indicated by the pod status in the exhibit. When a container exceeds its memory limit, Kubernetes terminates it with an OOMKilled exit code. Increasing the memory limit in the pod spec allows the container to allocate more memory, resolving the failure.

Exam trap

CompTIA often tests the distinction between resource exhaustion errors (OOMKilled vs. CPU throttling) and configuration errors (driver issues), leading candidates to incorrectly attribute a memory limit issue to a hardware or driver problem.

How to eliminate wrong answers

Option A is wrong because the exhibit shows no CPU-related errors or resource pressure; the failure is due to memory exhaustion, not insufficient CPU. Option B is wrong because GPU driver issues would manifest as device plugin errors or initialization failures, not an OOMKilled status. Option C is wrong because the model size is not directly indicated as the cause; the pod is failing due to memory limits, and using a smaller model might reduce memory usage but does not address the misconfigured resource limit.

136
MCQmedium

A data scientist is building a recommendation system using Apache Spark for feature engineering. They need to process streaming user click data in real-time before feeding into the model. Which tool should they use for the streaming data ingestion?

A.Amazon S3
B.Apache Kafka
C.Airflow
D.Snowflake
AnswerB

Kafka provides a distributed, partitioned, replicated commit log that ingests high-throughput click streams with low latency and durable buffering, decoupling producers from Spark Structured Streaming consumers. This satisfies the requirement to process real-time streaming click data before feature engineering.

Why this answer

Apache Kafka is the correct choice because it is a distributed streaming platform designed for high-throughput, fault-tolerant, real-time data ingestion. It acts as a durable message broker that can ingest streaming click data and make it available for Spark Structured Streaming to process in micro-batches or continuous processing mode, which is essential for real-time feature engineering in a recommendation system.

Exam trap

CompTIA often tests the distinction between storage, orchestration, and streaming tools, and the trap here is that candidates confuse batch-oriented tools like S3 or Airflow with real-time streaming ingestion, overlooking Kafka's role as a dedicated event streaming platform.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is an object storage service, not a streaming ingestion tool; it lacks the low-latency, pub-sub messaging capabilities required for real-time data streaming. Option C is wrong because Airflow is a workflow orchestration tool for scheduling batch jobs, not a real-time streaming ingestion platform; it cannot handle continuous, event-driven data streams. Option D is wrong because Snowflake is a cloud-based data warehouse optimized for analytical queries on structured data, not for real-time streaming ingestion; it does not provide a pub-sub or message queue interface for live click data.

137
MCQhard

A cybersecurity firm is developing an AI system to detect zero-day malware using behavior analysis. The team collects a dataset of 1,000 malware samples and 10,000 benign files from corporate endpoints. The model is a random forest classifier. After deployment, the false positive rate is 5%, which is acceptable, but the detection rate for new malware variants drops to 30%. The security analyst suspects the model is overfitting to the specific malware families in the training set. Which improvement should the team implement first?

A.Use a boosting ensemble instead of bagging
B.Collect more malware samples from the same families
C.Replace the random forest with a deep neural network
D.Engineer features that capture generic behavioral patterns
AnswerD

Generic behavioural features let the classifier generalise to unseen malware families instead of memorising training-set signatures. This directly addresses the overfitting that caused detection of new variants to fall to 30%, improving generalisation before other changes.

Why this answer

The core issue is that the model has overfitted to the specific malware families in the training set, causing poor generalization to unseen zero-day variants. Engineering features that capture generic behavioral patterns (e.g., API call sequences, file system interactions, network connection anomalies) reduces reliance on family-specific signatures, improving detection of novel malware. This directly addresses the root cause of the 30% detection rate drop without introducing new model complexity or data imbalance issues.

Exam trap

CompTIA often tests the misconception that more complex models (boosting, DNNs) automatically improve performance, when in reality, feature engineering to address the specific failure mode (overfitting to training families) is the most effective first step.

How to eliminate wrong answers

Option A is wrong because boosting ensembles (e.g., AdaBoost, XGBoost) are more prone to overfitting on noisy data than bagging (Random Forest), which would exacerbate the existing overfitting problem. Option B is wrong because collecting more samples from the same families reinforces the model's bias toward those specific patterns, worsening generalization to new variants. Option C is wrong because replacing Random Forest with a deep neural network (DNN) typically requires significantly more data to avoid overfitting, and with only 1,000 malware samples, a DNN would likely perform worse, not better.

138
MCQeasy

During data preparation for a classification model, the data scientist notices that one class has 95% of the samples and the other has only 5%. Which technique is MOST appropriate to address this imbalance?

A.Shuffle the data randomly before each training epoch
B.Remove the minority class samples entirely
C.Use a larger learning rate to force the model to pay attention to the minority class
D.Apply SMOTE (Synthetic Minority Over-sampling Technique) to generate synthetic samples for the minority class
AnswerD

SMOTE interpolates new minority-class points between existing minority neighbours, directly correcting the 95/5 skew. Unlike random oversampling, it does not merely duplicate rows, reducing overfitting risk. This addresses the stated class imbalance before training the classifier.

Why this answer

SMOTE generates synthetic samples of the minority class by interpolating between existing minority instances in feature space, which balances the class distribution and prevents the model from ignoring the minority class. With a 95/5 split, standard classifiers tend to predict the majority class almost exclusively, achieving high accuracy but poor recall on the minority class. SMOTE is the standard, well-established technique for this exact scenario.

Exam trap

AI0-001 often tests whether candidates confuse data-level techniques (SMOTE, resampling) with algorithm-level techniques (class weights, focal loss) — a larger learning rate is a tempting but incorrect 'make the model care' answer.

How to eliminate wrong answers

Option A is wrong because shuffling data before each epoch only changes sample order and does not alter class distribution, so the imbalance persists. Option B is wrong because removing minority samples eliminates the class entirely, making the problem worse and destroying the model's ability to learn it. Option C is wrong because a larger learning rate affects optimization step size, not class weighting, and can cause instability or divergence rather than fixing imbalance.

139
MCQmedium

A data engineer needs to design a data pipeline for a real-time fraud detection system. The system requires low-latency processing of streaming transactions. Which architecture is most appropriate?

A.Stream processing with Apache Kafka and Flink
B.Data lake with Apache Spark
C.Batch processing with Apache Hadoop
D.Microservices architecture with REST APIs
AnswerA

Kafka ingests high-throughput transaction streams durably, while Flink performs stateful, event-time windowed computation with millisecond latency, satisfying the low-latency streaming requirement. Batch architectures such as Hadoop or Spark microbatching introduce latency unsuited to real-time fraud detection.

Why this answer

Apache Kafka provides a distributed, fault-tolerant event streaming platform that ingests high-throughput transaction data with low latency, while Apache Flink offers true stream processing with exactly-once semantics and sub-second event-time processing. Together, they enable real-time fraud detection by analyzing transactions as they arrive, without the delays inherent in batch or micro-batch approaches.

Exam trap

CompTIA often tests the distinction between true stream processing (e.g., Flink, Kafka Streams) and micro-batch or near-real-time processing (e.g., Spark Streaming), where candidates mistakenly assume that any 'streaming' API (like Spark Streaming) is equivalent to low-latency stream processing.

How to eliminate wrong answers

Option B is wrong because a data lake with Apache Spark typically relies on micro-batch processing (e.g., Spark Streaming with a minimum batch interval of ~100ms), which introduces higher latency than true stream processing and is unsuitable for sub-second fraud detection. Option C is wrong because batch processing with Apache Hadoop (e.g., MapReduce) is designed for high-throughput, high-latency processing of large static datasets, not for real-time streaming where transactions must be evaluated within milliseconds. Option D is wrong because microservices architecture with REST APIs is a design pattern for building distributed services, not a data pipeline technology; REST APIs introduce synchronous request-response overhead and cannot natively handle continuous, unbounded data streams with low-latency stateful processing.

140
MCQmedium

A retail company's ML platform team notices that one of their production models has begun returning predictions with a drastically different distribution than during training. The monitoring dashboard shows the input feature distributions have shifted but no code or model artifacts have changed. The team wants to automatically trigger a retraining pipeline when this condition is detected. Which approach should they implement?

A.Configure a data drift monitor that computes statistical distance between live inference inputs and the training baseline, and wire its alert to the retraining pipeline trigger.
B.Enable concept drift detection by comparing the relationship between features and labels over time and trigger retraining when the mapping changes.
C.Implement a model versioning system that records the exact training data hash and triggers retraining whenever the hash of the incoming data differs from the stored hash.
D.Set up a model performance monitor that tracks prediction accuracy against ground truth labels and triggers retraining when accuracy drops below a threshold.
AnswerA

This is correct because the scenario describes a change in input feature distributions with no change to model code or artifacts, which is data drift. A drift monitor comparing live inputs to the training baseline detects this and can trigger retraining automatically.

Why this answer

The scenario describes data drift: input feature distributions have changed while the model and code remain the same. A data drift monitor that compares live inputs to the training baseline is the correct tool to detect this and can be integrated with the retraining pipeline. Other monitoring types either require labels, focus on label relationships, or are overly sensitive.

Exam trap

The trap here is confusing data drift with concept drift or model performance degradation, leading to selection of a monitor that does not directly detect input distribution changes.

141
MCQeasy

An ML team wants to prevent attackers from stealing a proprietary model by repeatedly querying the public API. Which defense is most effective?

A.Using a smaller model to reduce query cost
B.Encrypting model weights at rest
C.Adding random noise to all outputs
D.Rate limiting on the API endpoint
AnswerD

Rate limiting caps the number of queries a client can make in a given window, which directly throttles the high-volume repeated querying that model-extraction attacks depend on. By restricting query throughput, attackers cannot gather enough input-output pairs to reconstruct the proprietary model, satisfying the stem's requirement to prevent theft via the public API.

Why this answer

Rate limiting restricts the number of API requests a single client can make within a given time window, directly impeding an attacker's ability to collect enough query-response pairs to reconstruct or steal the model. This defense targets the attack vector itself—repeated queries—without degrading model performance for legitimate users. Techniques like token bucket or sliding window rate limiting are commonly implemented at the API gateway level.

Exam trap

CompTIA often tests the misconception that encryption or obfuscation of model artifacts is sufficient to prevent extraction attacks, when in fact the primary threat is from live API queries that bypass those protections.

How to eliminate wrong answers

Option A is wrong because using a smaller model reduces computational cost but does not prevent an attacker from querying the API repeatedly to extract the model's behavior; the attack surface remains unchanged. Option B is wrong because encrypting model weights at rest protects against offline theft of stored model files, but does nothing to stop an attacker from querying the live API endpoint to perform model extraction. Option C is wrong because adding random noise to all outputs degrades the model's accuracy for all users and can be mitigated by averaging multiple queries, making it an ineffective and impractical defense against model stealing.

142
Multi-Selecteasy

A machine learning engineer is preparing to train a deep neural network for image classification. To avoid overfitting, which TWO techniques should the engineer apply? (Select TWO.)

Select 2 answers
A.Use dropout regularization.
B.Use data augmentation.
C.Increase the number of layers.
D.Remove all non-linear activation functions.
E.Reduce the training dataset size.
AnswersA, B

Dropout is a regularization technique that helps prevent overfitting by randomly dropping units.

Why this answer

Dropout regularization is a technique that randomly drops a fraction of neurons during training, which prevents the network from relying too heavily on any single neuron and reduces co-adaptation. This acts as a form of ensemble learning and significantly reduces overfitting by improving generalization.

Exam trap

The CompTIA AI+ exam often tests the misconception that increasing model complexity (like adding layers) or reducing data helps with overfitting, when in reality these actions worsen it, while regularization and data augmentation are the correct countermeasures.

143
MCQhard

A cybersecurity firm is building an anomaly detection system for network traffic. The dataset contains millions of connection records with dozens of features, but only 0.1% are labeled as malicious. The team needs a model that can flag suspicious connections while minimizing false positives that overwhelm analysts. Which approach is most appropriate?

A.Apply principal component analysis (PCA) to reduce dimensionality, then use a threshold on reconstruction error to detect anomalies.
B.Use k-means clustering with k set to the number of known attack types and treat points far from centroids as anomalies.
C.Train an isolation forest or autoencoder on the normal traffic to learn its structure, then flag deviations as anomalies.
D.Train a supervised gradient boosting classifier on the labeled data and use class weights to handle imbalance.
AnswerC

With only 0.1% malicious labels, supervised classification is impractical due to extreme class imbalance. Unsupervised anomaly detection methods like isolation forest and autoencoders learn the distribution of normal traffic and flag deviations, requiring no labels. This directly suits the scenario, and by tuning the contamination parameter or reconstruction error threshold, the team can control the false positive rate to keep analysts from being overwhelmed.

Why this answer

Extreme class imbalance with only 0.1% malicious labels makes supervised learning unreliable. Unsupervised anomaly detection methods like isolation forest and autoencoders learn the structure of normal traffic without requiring labels, then flag deviations. Isolation forest isolates anomalies by random partitioning, while autoencoders flag high reconstruction error.

Both allow threshold tuning to balance detection and false positives, which is critical for analyst workload. This makes them the most appropriate choice for the cybersecurity scenario.

Exam trap

The trap here is assuming that class weighting or dimensionality reduction alone can overcome a 0.1% positive rate, when the scarcity of labels makes unsupervised anomaly detection the only practical path.

144
MCQhard

A machine learning engineer is troubleshooting a recurrent neural network that fails to learn long-range dependencies in sequential data. The gradients are computed using backpropagation through time. Which phenomenon is most likely occurring, and what architectural change would best address it?

A.Underfitting; increase the number of time steps
B.Vanishing gradients; use LSTM or GRU units
C.Exploding gradients; apply gradient clipping
D.Overfitting; reduce the number of layers
AnswerB

Backpropagation through time multiplies many small Jacobian terms, so gradients shrink exponentially across long sequences, preventing the network from learning distant dependencies. LSTM and GRU units introduce gated cell states that carry information along a near-constant error path, preserving gradient magnitude over long ranges.

Why this answer

In standard RNNs, backpropagation through time (BPTT) multiplies gradients across many time steps, causing them to shrink exponentially (vanishing gradients). This prevents the network from learning long-range dependencies. LSTM or GRU units introduce gating mechanisms that preserve gradient flow over many time steps, directly solving this problem.

Exam trap

The trap here is that candidates may confuse 'failure to learn long-range dependencies' with exploding gradients, but the correct clue is the inability to capture distant patterns, not training instability or NaN losses.

How to eliminate wrong answers

Option A is wrong because increasing the number of time steps would exacerbate the vanishing gradient problem, not fix it; underfitting is not the core issue here. Option C is wrong because exploding gradients cause large gradient values and training instability, but the described symptom (failure to learn long-range dependencies) is classic vanishing gradients, not exploding gradients; gradient clipping addresses exploding gradients, not vanishing ones. Option D is wrong because overfitting is characterized by high variance and poor generalization, not by an inability to learn long-range patterns; reducing layers would reduce capacity and could worsen underfitting, not solve the gradient propagation issue.

145
MCQmedium

An AI system used for hiring has been found to exhibit racial bias against certain candidates. Which step should the organization take to mitigate this?

A.Remove all demographic features from the model.
B.Use a different algorithm that is inherently unbiased.
C.Regularly audit model predictions across demographic groups and retrain with fairness constraints.
D.Hire more diverse data scientists.
AnswerC

Auditing predictions across demographic groups exposes disparate impact that aggregate accuracy hides, satisfying the need to detect racial bias. Retraining with fairness constraints then adjusts the model's decision boundary to reduce that measured disparity, rather than merely documenting it. This directly targets the biased hiring outcomes described.

Why this answer

Bias in AI systems is often embedded in training data or model behavior, not just in feature selection. Regularly auditing predictions across demographic groups and retraining with fairness constraints (e.g., demographic parity or equalized odds) allows the organization to detect and correct disparate impact without sacrificing model performance. This aligns with the AI0-001 focus on continuous monitoring and iterative improvement in AI operations.

Exam trap

CompTIA often tests the misconception that removing sensitive attributes (like race or gender) automatically makes a model fair, when in reality proxy features and biased training data can perpetuate discrimination.

How to eliminate wrong answers

Option A is wrong because simply removing demographic features does not eliminate bias; proxy features (e.g., zip code, education level) can still encode the same discriminatory patterns, and the model may learn biased correlations from the remaining data. Option B is wrong because no algorithm is inherently unbiased; bias arises from data, labeling, and deployment context, so switching algorithms without addressing root causes will not guarantee fairness. Option D is wrong because hiring more diverse data scientists, while beneficial for broader perspectives, does not directly mitigate existing model bias; technical interventions like auditing and retraining with fairness constraints are required.

146
MCQhard

A machine learning engineer is preparing a dataset for a natural language processing task. The dataset contains text reviews with varying lengths, and the engineer plans to use a transformer model. Which preprocessing step is most critical to ensure the model can handle the input effectively?

A.Convert all text to lowercase and remove punctuation to reduce vocabulary size.
B.Perform stemming or lemmatization to reduce words to their base forms.
C.Tokenize the text into subword units and pad or truncate sequences to a fixed maximum length.
D.Apply one-hot encoding to each word in the vocabulary to create binary vectors.
AnswerC

Transformer models require input sequences of uniform length within a batch. Tokenization into subword units (e.g., WordPiece or BPE) handles out-of-vocabulary words and reduces vocabulary size. Padding shorter sequences and truncating longer ones to a fixed maximum length ensures batch processing. This is essential for efficient training and inference with transformers.

Why this answer

Transformer models require fixed-length input sequences, so tokenizing into subword units and padding/truncating to a uniform length is essential. This enables batch processing and handles out-of-vocabulary words. Other steps like lowercasing, one-hot encoding, or stemming are either not critical or incompatible with transformer input expectations.

Exam trap

The trap here is focusing on traditional text normalization steps like lowercasing or stemming while overlooking the transformer-specific need for uniform sequence lengths via tokenization and padding.

147
Multi-Selecthard

An organization is implementing an AI governance framework. Which THREE components are essential for compliance with ethical AI standards?

Select 3 answers
A.Data privacy protection measures (e.g., differential privacy).
B.Open-source licensing of all models.
C.Maximizing model accuracy to increase revenue.
D.Model explainability and interpretability mechanisms.
E.Regular bias auditing of models.
AnswersA, D, E

Differential privacy adds calibrated noise so individual records cannot be re-identified from model outputs or training data, directly satisfying the ethical standard's data privacy protection requirement. It is a technical control, not merely a policy statement, making it essential within an enforceable AI governance framework.

Why this answer

Option A (data privacy protection measures such as differential privacy) is essential because ethical AI standards require safeguarding personal data, and techniques like differential privacy provide formal, quantifiable guarantees that individual records cannot be inferred from model outputs, supporting regulations like GDPR. Option D (model explainability and interpretability mechanisms) is essential because stakeholders must understand how decisions are reached; methods such as SHAP, LIME, or inherently interpretable models enable accountability and contestability required by ethical frameworks. Option E (regular bias auditing of models) is essential because systematic auditing detects disparate impact across protected groups, using metrics like demographic parity or equalized odds, ensuring fairness is continuously monitored rather than assumed.

Option B is not required because ethical compliance concerns how models are governed and used, not whether they are open-sourced; proprietary models can be fully ethical. Option C is incorrect because maximizing accuracy for revenue is a business objective, not an ethical compliance requirement, and can even conflict with fairness and privacy goals.

Exam trap

The AI0-001 exam often tests the misconception that open-source licensing or maximizing accuracy are ethical imperatives, when in fact they are operational or business choices that do not directly satisfy the core pillars of ethical AI (privacy, fairness, transparency, accountability).

148
MCQeasy

Which similarity metric is MOST appropriate for comparing dense vector embeddings in a vector store used for document retrieval, when the embeddings are normalized to unit length?

A.Jaccard similarity
B.Manhattan distance
C.Cosine similarity
D.Euclidean distance
AnswerC

With unit-length embeddings, cosine similarity and dot product give identical rankings, but cosine similarity directly measures the angle between vectors, remaining invariant to magnitude. It satisfies the stem's normalisation constraint and is the standard metric for dense retrieval in vector stores.

Why this answer

Cosine similarity measures the angle between two vectors and is the standard metric for comparing dense embeddings in vector stores. When embeddings are normalized to unit length, cosine similarity is mathematically equivalent to the dot product, making it both efficient and semantically meaningful for document retrieval. It focuses on orientation rather than magnitude, which aligns with how embedding models encode semantic similarity.

Exam trap

AI0-001 often tests the relationship between normalization and similarity metrics — candidates may pick Euclidean distance thinking it is equivalent, but cosine similarity is the conventional and mathematically clean choice for unit-length embeddings.

How to eliminate wrong answers

Option A is wrong because Jaccard similarity compares set overlap (intersection over union) and is used for binary or categorical features, not dense continuous vectors. Option B is wrong because Manhattan distance (L1) measures absolute coordinate differences and is sensitive to vector magnitude, making it less appropriate for normalized embeddings where direction matters more than magnitude. Option D is wrong because Euclidean distance (L2) on normalized vectors is monotonically related to cosine similarity but is less commonly used as the primary retrieval metric and can be less numerically stable in high dimensions; cosine similarity is the conventional choice for embedding stores.

149
MCQeasy

A data scientist trains a linear regression model on housing prices. The training error is low, but test error is high. What is the most likely issue?

A.Overfitting
B.Multicollinearity
C.Data leakage
D.Underfitting
AnswerA

Overfitting occurs when the model captures noise and training-specific patterns, fitting training data closely while generalising poorly. The gap between low training error and high test error is the defining symptom, indicating the model has learned idiosyncrasies rather than the underlying relationship.

Why this answer

Low training error combined with high test error is the classic symptom of overfitting, where the model has memorized the training data, including its noise, rather than learning the underlying patterns. This causes the model to perform poorly on unseen data, which is exactly what the high test error indicates.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by presenting a scenario where training error is low but test error is high, leading candidates to mistakenly choose underfitting because they focus only on the high test error without considering the low training error.

How to eliminate wrong answers

Option B (Multicollinearity) is wrong because while it can inflate the variance of coefficient estimates and make them unstable, it does not typically cause a large gap between low training error and high test error; the model can still fit the training data well. Option C (Data leakage) is wrong because data leakage usually results in overly optimistic performance on both training and test sets during evaluation, not a high test error after low training error. Option D (Underfitting) is wrong because underfitting produces high error on both the training and test sets, not low training error with high test error.

150
MCQhard

An e-commerce company deploys a recommendation system using collaborative filtering. After launch, the system shows high accuracy for popular items but fails to recommend niche products to users who would likely buy them. Which technique should the team implement to improve recommendations for long-tail items?

A.Apply matrix factorization with higher latent factors
B.Switch to a hybrid filtering approach that incorporates item metadata
C.Increase the weight of popular items in the recommendation score
D.Collect more user interaction data over time
AnswerB

Collaborative filtering relies on user-item interaction overlap, so sparse long-tail items receive poor representations. A hybrid approach adds item metadata (content features), letting the model recommend niche products even when interaction data is scarce, directly addressing the stem's long-tail failure.

Why this answer

Collaborative filtering relies on user-item interactions, which are sparse for niche products (the long tail). A hybrid filtering approach that incorporates item metadata (e.g., category, description, attributes) can bridge the gap by using content-based signals to recommend niche items even when interaction data is limited. This directly addresses the cold-start and sparsity problems for long-tail items.

Exam trap

CompTIA often tests the misconception that more data or higher model complexity (like more latent factors) automatically solves sparsity, when in fact the core issue is the lack of interaction signals for niche items, which requires a hybrid approach to incorporate auxiliary information.

How to eliminate wrong answers

Option A is wrong because increasing latent factors in matrix factorization can lead to overfitting and does not inherently solve the sparsity problem for long-tail items; it may even amplify noise. Option C is wrong because increasing the weight of popular items would further bias recommendations toward the head of the distribution, worsening the neglect of niche products. Option D is wrong because simply collecting more user interaction data over time does not guarantee that long-tail items will receive sufficient interactions; the data will still be skewed toward popular items, and the system needs a mechanism to leverage non-interaction signals like metadata.

Page 1

Page 2 of 13

Page 3