Courseiva

CompTIA AI+ AI0-001 (AI0-001) — Questions 526–600

962 questions total · 13pages · All types, answers revealed

Page 7

Page 8 of 13

Page 9
526
Multi-Selecteasy

A data scientist is monitoring a deployed image classification model. Which TWO actions are best practices for detecting model drift? (Choose 2.)

Select 2 answers
A.Schedule automatic weekly retraining of the model.
B.Increase the model's complexity to improve generalization.
C.Use a holdout test set to periodically evaluate model accuracy.
D.Monitor the average prediction confidence of the model.
E.Track the distribution of input data over time.
AnswersC, E

A holdout test set provides labelled ground truth, letting the team measure accuracy decay against known-correct outputs over time. This detects concept drift that unlabelled input monitoring alone cannot confirm, satisfying the requirement to identify genuine performance degradation.

Why this answer

Option C is correct because periodically scoring the model against a fixed holdout test set gives a direct, repeatable measurement of accuracy degradation, which is the core signal of model drift. Option E is correct because tracking the distribution of input features over time (e.g., via histograms or statistical tests like KL divergence or PSI) detects data drift, which typically precedes and causes model drift. Option A is not a detection practice but a remediation action, and retraining blindly on a schedule does not identify drift.

Option B is a modeling change that may affect generalization but does nothing to detect drift in a deployed model. Option D is weaker than C and E because average prediction confidence can shift for reasons unrelated to drift and is not a reliable standalone drift indicator.

Exam trap

CompTIA often tests the distinction between detection and remediation actions, so candidates mistakenly choose retraining (Option A) as a detection method when it is actually a corrective action.

527
MCQeasy

A retail company runs a customer-facing chatbot backed by a large language model. The chatbot has access to a tool that looks up order status by order ID. A penetration tester finds that by typing a crafted sentence, a user can make the chatbot call the order-status tool with an arbitrary order ID belonging to another customer and read the response. Which control most directly prevents this unauthorized tool invocation?

A.Enforce authorization checks in the tool backend so it only returns orders belonging to the authenticated user.
B.Add a system prompt instructing the model never to reveal other customers' order details.
C.Increase the model's temperature setting to make its responses less predictable.
D.Log every tool call the chatbot makes and review the logs daily for suspicious order IDs.
AnswerA

The vulnerability is that the tool trusts the order ID supplied through the model without verifying ownership. Moving the authorization decision into the tool backend, where it can compare the requested order against the authenticated session's customer ID, ensures that even a manipulated prompt cannot retrieve another customer's data. This is the most direct fix because it removes reliance on the model to enforce access control.

Why this answer

The flaw is that the tool performs a lookup based solely on a model-supplied order ID, with no check that the order belongs to the authenticated user. Enforcing authorization in the tool backend removes trust from the model and blocks cross-customer access regardless of how the prompt is crafted. Prompt instructions, temperature changes, and after-the-fact logging do not establish that access boundary.

Exam trap

The trap here is trusting the language model to enforce access control through instructions, when authorization must live in the tool backend outside the model's control.

528
Multi-Selecthard

A security engineer is hardening an LLM application against indirect prompt injection attacks. Which TWO controls are MOST effective? (Select two.)

Select 2 answers
A.Output filtering
B.Input validation and sanitization
C.Rate limiting
D.Differential privacy
E.Federated learning
AnswersA, B

Output filtering inspects the model's response before it reaches users or downstream systems, catching injected instructions that survived input handling. It provides defence in depth against indirect injection, where malicious content arrives via retrieved documents or tool output rather than direct user input.

Why this answer

Output filtering (A) is correct because it inspects the model's generated response before it reaches the user or downstream system, allowing detection and blocking of leaked system prompts, injected instructions, or malicious content that survived the attack chain. Input validation and sanitization (B) is correct because indirect prompt injection arrives through untrusted external content (retrieved documents, web pages, tool outputs), so stripping or neutralizing embedded instructions, delimiters, and control tokens before they enter the context window reduces the attack surface. Rate limiting (C) only throttles request volume and does not address the semantic content of injected prompts.

Differential privacy (D) protects training-data privacy by adding noise, not runtime prompt integrity. Federated learning (E) is a distributed training paradigm and has no bearing on inference-time injection defense.

Exam trap

AI0-001 often tests the misconception that generic security controls like rate limiting or privacy techniques like differential privacy defend against prompt injection, when only input validation and output filtering address the attack path.

529
MCQhard

A large hospital system deploys an AI triage system for emergency rooms. The system uses patient vitals and symptoms to recommend treatment priority. Six months after deployment, complaints arise that the system frequently underestimates the severity of symptoms for patients from certain ethnic backgrounds. A data scientist runs a bias audit and finds that the model's false negative rate is 20% higher for the minority group. The hospital's AI governance board requires immediate corrective action. The data science team has limited resources and cannot retrain the entire model from scratch. They have access to the training data, which is imbalanced. The model is a gradient boosted tree. Which course of action best addresses the bias while minimizing operational impact?

A.Rebalance the training data using SMOTE and retrain the model
B.Use adversarial debiasing during training to remove protected attribute correlations
C.Post-process the model's predictions by adjusting thresholds for the minority group
D.Replace the model with a simpler logistic regression model to improve interpretability
AnswerC

Threshold adjustment is a post-processing intervention applied at inference, so it corrects the disparate false negative rate without retraining the gradient boosted tree. This satisfies the constraint of limited resources and minimal operational impact, equalising error rates across groups while leaving the existing model intact.

Why this answer

Post-processing by adjusting decision thresholds for the minority group directly compensates for the higher false negative rate without requiring retraining. Since the team has limited resources and cannot retrain the entire gradient boosted tree model, this approach minimizes operational impact while addressing the bias. The threshold adjustment effectively lowers the probability cutoff for the minority group, making the model more sensitive to their symptoms and reducing underestimation of severity.

Exam trap

CompTIA often tests the misconception that bias mitigation always requires retraining or complex algorithmic changes, when in fact post-processing threshold adjustments can be a quick, effective fix for deployed models with limited resources.

How to eliminate wrong answers

Option A is wrong because SMOTE rebalances the training data by oversampling the minority class, but retraining the entire gradient boosted tree model from scratch is resource-intensive and contradicts the constraint of limited resources; moreover, SMOTE may introduce synthetic noise that degrades model performance. Option B is wrong because adversarial debiasing is a training-time technique that requires modifying the model architecture and retraining, which is not feasible given the limited resources and the fact that the model is already deployed; it also does not directly address the false negative rate disparity without full retraining. Option D is wrong because replacing the model with a simpler logistic regression model would require retraining and likely reduce predictive performance, especially for complex interactions in patient vitals and symptoms, and does not guarantee bias reduction; interpretability alone does not correct the existing bias.

530
MCQmedium

A data scientist is building a model to predict the likelihood of a customer defaulting on a loan. The dataset contains a feature 'debt_to_income_ratio' that is highly skewed with a long right tail. The scientist decides to apply a logarithmic transformation to this feature. Which statement best describes the effect of this transformation?

A.It normalizes the feature to a range between 0 and 1, which is required for all machine learning algorithms.
B.It converts the feature into a categorical variable, which is useful for tree-based models.
C.It reduces the impact of outliers and makes the distribution more symmetric, which can help linear models perform better.
D.It increases the variance of the feature, allowing the model to capture more complex patterns.
AnswerC

A logarithmic transformation compresses the scale of large values, pulling in the right tail and making the distribution more normal-like. This reduces the influence of extreme outliers and can improve the performance of linear models that assume normally distributed errors or linear relationships.

Why this answer

Log transformation is commonly used to handle right-skewed data. It reduces skewness, stabilizes variance, and makes the distribution more symmetric, which can improve the performance of linear models and neural networks that are sensitive to feature scaling. It does not increase variance, convert to categorical, or normalize to a fixed range.

Exam trap

The trap here is confusing log transformation with normalization or standardization, leading to the incorrect belief that it scales features to a specific range.

531
MCQhard

An AI model is deployed to a mobile app with limited computational resources. The model is a deep neural network with high latency. Which technique is best to reduce inference time?

A.Increase batch size
B.Add more layers
C.Use a larger model
D.Quantization
AnswerD

Quantization converts weights and activations from 32-bit floats to 8-bit integers, shrinking memory footprint and enabling faster integer arithmetic on resource-constrained mobile hardware. This directly cuts inference latency while preserving acceptable accuracy, satisfying the limited-compute constraint.

Why this answer

Quantization reduces the precision of the model's weights and activations (e.g., from 32-bit floating point to 8-bit integer), which decreases memory footprint and speeds up computation on resource-constrained devices like mobile phones. This directly lowers inference latency without requiring additional hardware or architectural changes.

Exam trap

The AI0-001 exam often tests the misconception that increasing batch size or model size improves performance on edge devices, when in fact these techniques increase resource demands and latency in low-resource environments.

How to eliminate wrong answers

Option A is wrong because increasing batch size improves throughput (samples per second) but does not reduce per-sample latency; it actually increases memory usage and can worsen latency on mobile devices with limited resources. Option B is wrong because adding more layers increases the model depth, which increases computational complexity and latency, making inference slower. Option C is wrong because using a larger model (more parameters) increases both memory and compute requirements, directly increasing inference time on constrained devices.

532
MCQmedium

An AI model's performance drops significantly in production compared to testing. The data shows distribution shift. What is the best first step?

A.Add more features
B.Retrain model with new data
C.Use a different algorithm
D.Reduce model complexity
AnswerB

Retraining on recent data is the first practical step because distribution shift means the learned mapping no longer matches current input statistics; refreshing training data realigns the model with production. Monitoring alone would not restore performance, and the stem already identifies shift as the cause.

Why this answer

When distribution shift occurs, the model's assumptions about the data distribution no longer hold, so the best first step is to retrain the model with new data that reflects the current distribution. This directly addresses the root cause by updating the model to learn the new patterns, rather than applying superficial fixes like adding features or changing algorithms.

Exam trap

AI0-001 often tests the tendency to jump to algorithmic solutions (different algorithm, more features) when the root cause is data distribution, and the correct first step is usually data-centric.

How to eliminate wrong answers

Option A is wrong because adding more features does not address distribution shift and may introduce noise or overfitting. Option C is wrong because using a different algorithm does not fix the underlying data distribution mismatch and may perform worse without proper retraining. Option D is wrong because reducing model complexity may help with overfitting but does not resolve distribution shift, which requires adapting to new data.

533
Multi-Selectmedium

A data scientist is using differential privacy to protect individual privacy in a training dataset. Which TWO actions are correct implementations of differential privacy?

Select 2 answers
A.Train the model on a small subset of data to reduce exposure
B.Remove all personally identifiable information (PII) from the dataset
C.Aggregate data into groups before training
D.Set a privacy budget (epsilon) to limit information leakage
E.Add noise to the training data to mask individual contributions
AnswersD, E

The privacy budget epsilon quantifies the privacy guarantee and is a core concept of differential privacy.

Why this answer

Setting a privacy budget (epsilon) is a core mechanism in differential privacy that quantifies and limits the amount of information leaked about any individual in the dataset. By controlling epsilon, the data scientist can formally bound the privacy loss, ensuring that the model's outputs do not reveal whether any specific individual's data was included in training.

Exam trap

A common pitfall in this question is thinking that removing PII or aggregating data is sufficient for differential privacy. In reality, differential privacy requires a formal mathematical framework with noise addition and a privacy budget parameter. CompTIA often tests this distinction.

534
MCQmedium

A team is developing an AI agent to assist users with multi-step tasks such as booking a flight, reserving a hotel, and scheduling a car rental. The agent needs to reason about the order of steps and handle dependencies. Which pattern is BEST suited?

A.Simple tool use without reasoning
B.Using a single prompt with all instructions
C.Fine-tuning a model to output all steps at once
D.ReAct pattern (Reasoning and Acting)
AnswerD

ReAct interleaves reasoning traces with tool-calling actions, letting the agent decide each next step based on observed results. This supports sequencing dependent tasks such as flight, hotel, and car rental bookings, where later steps rely on earlier outcomes.

Why this answer

The ReAct pattern (Reasoning and Acting) is best suited because it interleaves reasoning traces with tool calls, allowing the agent to dynamically plan and adjust steps based on intermediate results. For multi-step tasks with dependencies (e.g., booking a flight before a hotel), ReAct enables the agent to reason about order, handle failures, and call external APIs step-by-step, which is essential for robust task completion.

Exam trap

The key trap is that candidates overlook the need for dynamic reasoning and tool interaction, assuming that a single large prompt or fine-tuned output can handle all multi-step tasks. However, only the ReAct pattern provides the necessary step-by-step planning and tool use.

How to eliminate wrong answers

Option A is wrong because simple tool use without reasoning lacks the ability to plan or handle dependencies; it can only execute isolated function calls without context. Option B is wrong because using a single prompt with all instructions cannot adapt to dynamic changes or intermediate results; it assumes a static plan that fails if any step requires conditional logic or error recovery. Option C is wrong because fine-tuning a model to output all steps at once (single-shot generation) cannot handle real-time feedback from external systems or adapt to variable execution order, making it brittle for interactive multi-step workflows.

535
MCQhard

A hospital's radiology AI triage model was validated at 94% sensitivity on a curated research dataset. After six months in production, clinicians report that it misses many positive cases on images from a newly installed scanner. The data science team confirms the model has not been retrained. Which action should the team take FIRST to diagnose and correct the problem?

A.Increase the model's classification threshold so that more cases are flagged as positive.
B.Compare the statistical distribution of production images from the new scanner against the training data to detect covariate shift.
C.Fine-tune the model on the most recent production images that clinicians have labeled as missed positives.
D.Replace the triage model with a larger architecture trained from scratch on a combined dataset of all available scanners.
AnswerB

The model was validated on one data distribution and now receives images from different hardware, which strongly suggests covariate shift: the input features moved while the learned relationship stayed fixed. Measuring distribution differences such as intensity histograms, resolution, and noise characteristics identifies whether inputs are out of the training envelope, which is the root cause to establish before choosing any remediation.

Why this answer

Performance collapse on images from new hardware, with no retraining, points to a change in the input distribution rather than a change in the label relationship. Detecting and quantifying that shift through distribution comparison establishes the root cause before any intervention. Threshold changes, targeted fine-tuning, and full retraining are all remedies that presuppose a cause and can waste effort or introduce new bias if applied blindly.

Exam trap

The trap here is jumping to a remediation such as retraining or threshold tuning before confirming that the production inputs actually differ from the training distribution.

536
Multi-Selectmedium

A financial services firm runs a credit-scoring AI model in production. The compliance team wants continuous assurance that the deployed model still meets performance and fairness expectations as customer behavior changes. Which TWO operational practices best support this goal? (Choose two.)

Select 2 answers
A.Scheduled monitoring of key performance metrics such as AUC, precision, and recall against recent labeled outcomes
B.Regular fairness audits comparing approval rates and error rates across protected groups
C.Retraining the model on the full historical dataset every night regardless of any detected change
D.Storing all inference requests and responses in immutable logs for later forensic analysis
E.Increasing the model's parameter count to improve its capacity to fit customer data
AnswersA, B

Tracking performance metrics on recent labeled data detects degradation as customer behavior evolves, giving the firm evidence that the deployed model still performs as validated. Because labels for credit outcomes arrive with a delay, scheduling these evaluations ensures drift is caught before it materially harms decisions, which directly supports continuous assurance.

Why this answer

Continuous assurance requires active measurement of the deployed system on two fronts: predictive performance against fresh labeled outcomes and fairness across protected groups. These two practices detect degradation as conditions change and generate evidence for compliance. Retraining without evidence, enlarging the model, or merely archiving requests do not by themselves confirm that the model remains effective and equitable.

Exam trap

The trap here is selecting logging or retraining as assurance activities, when assurance specifically requires measuring performance and fairness outcomes over time.

537
Multi-Selecthard

During a security audit of an AI-powered code generation tool, the audit team discovers that the system prompt (which contains sensitive internal instructions) can be leaked through carefully crafted user inputs. Which THREE OWASP LLM Top 10 categories are MOST directly relevant to this finding?

Select 3 answers
A.Model denial of service
B.Prompt injection
C.Insecure output handling
D.Supply chain vulnerabilities
E.Sensitive information disclosure
AnswersB, C, E

Prompt injection directly enables this leak: crafted user inputs override or subvert the system prompt's instructions, causing the model to disclose its own sensitive contents. This satisfies the audit finding's core constraint — system prompt exposure via adversarial input — making it one of the three most relevant OWASP LLM Top 10 categories.

Why this answer

Option B (Prompt injection) is correct because the leak occurs through carefully crafted user inputs that manipulate the model into overriding or bypassing its system prompt instructions, which is the defining mechanism of prompt injection (direct or indirect). Option C (Insecure output handling) is correct because the system prompt and sensitive internal instructions are being emitted in the model's output without adequate validation, filtering, or sanitization before being returned to the user. Option E (Sensitive information disclosure) is correct because the core impact is the exposure of sensitive internal instructions contained in the system prompt, which is exactly the data-leakage concern this category covers.

Option A (Model denial of service) does not belong because there is no resource exhaustion, unbounded consumption, or availability impact described. Option D (Supply chain vulnerabilities) does not belong because the finding involves the model's own prompt handling and output, not compromised dependencies, models, or third-party components.

Exam trap

AI0-001 often tests the tendency to conflate the attack vector (Prompt Injection) with the resulting harm (Sensitive Information Disclosure) and the delivery flaw (Insecure Output Handling), causing candidates to select only one or two of the three when the question asks for all directly relevant categories.

538
MCQhard

A media company trains a video tagging model on a large dataset in the cloud. The model will run inference on-premises in a facility with intermittent network connectivity, and the operations team wants to avoid re-authoring the model for each target runtime. Which deployment artifact best meets these constraints?

A.A pickle file of the scikit-learn estimator
B.An ONNX model executed with the ONNX Runtime
C.The native TensorFlow SavedModel directory
D.A Docker image containing the full training environment
AnswerB

ONNX is an open, framework-neutral model representation, and ONNX Runtime provides a portable inference engine that runs on Windows, Linux, and edge devices without the original training framework. Exporting the trained model to ONNX lets the same artifact execute on-premises despite intermittent connectivity, and the team avoids re-authoring the model for each target runtime because the graph format is standardized.

Why this answer

ONNX defines a standardized computation graph, and ONNX Runtime executes that graph across platforms and hardware without requiring the original training framework. Exporting to ONNX gives the media company a single artifact that runs on-premises in a disconnected facility and removes the need to re-author the model for each inference runtime, satisfying both stated constraints.

Exam trap

The trap here is equating containerization with runtime portability, when a container still embeds one specific framework build and does not standardize the model graph.

539
Multi-Selectmedium

A data scientist is preparing a dataset for a binary classification model. The dataset has 1000 samples, with 800 positives and 200 negatives. To evaluate the model properly, which THREE steps should they take? (Select THREE)

Select 3 answers
A.Remove the minority class samples to make the dataset balanced
B.Use a stratified train-test split to preserve class proportions
C.Apply SMOTE (Synthetic Minority Over-sampling Technique) to balance the training set
D.Report only accuracy as the evaluation metric
E.Use precision, recall, and F1-score for evaluation
AnswersB, C, E

A stratified split preserves the 80:20 positive-to-negative ratio in both training and test sets, satisfying the requirement for proper evaluation. Without stratification, random splitting could yield test sets with skewed class proportions, distorting precision and recall.

Why this answer

Option B is correct because a stratified train-test split preserves the 80/20 class ratio (800 positives, 200 negatives) in both the training and test sets, ensuring the evaluation reflects the true class distribution and avoids sampling bias. Option C is correct because SMOTE generates synthetic minority-class samples by interpolating between existing minority instances, balancing the training set so the classifier does not become biased toward the majority class. Option E is correct because with imbalanced data, accuracy is misleading; precision, recall, and F1-score reveal how well the model identifies the minority class and balances false positives against false negatives.

Option A is wrong because deleting minority samples discards valuable information and worsens the class imbalance problem. Option D is wrong because a model predicting all positives would achieve 80% accuracy while completely failing on the minority class, making accuracy an unreliable metric here.

Exam trap

CompTIA often tests the misconception that removing minority samples or relying solely on accuracy is acceptable for imbalanced datasets, when in fact these approaches degrade model performance and evaluation validity.

540
MCQeasy

Which technique adds controlled noise to query results or training data to prevent an attacker from inferring whether a specific individual's data was included in the dataset?

A.Anonymisation
B.Federated learning
C.Differential privacy
D.Pseudonymisation
AnswerC

Differential privacy injects calibrated statistical noise into query outputs or training data, bounding how much any single record influences results. This mathematically limits inference about whether a specific individual's data was included, unlike anonymisation or masking.

Why this answer

Differential privacy is the technique that adds controlled statistical noise (e.g., Laplace or Gaussian noise) to query results or training data so that the presence or absence of any single individual's data cannot be inferred. It provides a mathematical guarantee (epsilon-differential privacy) that the output distribution is nearly identical whether or not a specific record is included. This directly matches the question's requirement.

Exam trap

AI0-001 often tests the confusion between anonymisation/pseudonymisation (which are data-masking techniques without formal guarantees) and differential privacy (which provides a mathematical privacy guarantee against inference), so candidates pick the familiar term instead of the rigorous one.

How to eliminate wrong answers

Option A is wrong because anonymisation removes or masks personally identifiable information but does not provide a formal guarantee against inference attacks — an attacker can still correlate quasi-identifiers to re-identify individuals (as shown in the Netflix Prize and AOL search log incidents). Option B is wrong because federated learning is a distributed training approach where models are trained locally on devices and only updates are aggregated; it protects raw data locality but does not by itself prevent inference about whether a specific individual's data was used. Option D is wrong because pseudonymisation replaces identifiers with pseudonyms but retains a mapping that can be reversed, so it does not prevent inference attacks and is not a formal privacy guarantee.

541
MCQeasy

A company wants to build a system that automatically tags uploaded images with objects they contain (e.g., 'car', 'tree', 'person'). Which AI application type is this?

A.Image classification/object detection
B.Recommendation system
C.Anomaly detection
D.Document intelligence
AnswerA

Object detection satisfies the requirement to tag multiple objects within one image, returning bounding boxes and labels per instance. Unlike image classification, which assigns a single label to the whole image, detection handles the stem's example of 'car', 'tree' and 'person' coexisting, so each object is identified separately.

Why this answer

The task of identifying and labeling objects (e.g., 'car', 'tree', 'person') within an image is a classic use case for image classification combined with object detection. Image classification assigns a single label to the entire image, while object detection localizes and classifies multiple objects within the image, which is exactly what the system requires.

Exam trap

The AI0-001 exam often tests the distinction between image classification (single label per image) and object detection (multiple localized objects), so candidates may mistakenly choose image classification alone when the question implies multiple objects per image.

How to eliminate wrong answers

Option B is wrong because recommendation systems analyze user behavior and preferences to suggest items (e.g., movies, products), not to identify objects in images. Option C is wrong because anomaly detection identifies unusual patterns or outliers in data (e.g., fraud detection), not the presence of common objects in images. Option D is wrong because document intelligence focuses on extracting text, structure, and information from documents (e.g., OCR, form processing), not on visual object recognition.

542
MCQeasy

A small logistics company wants to forecast next month's shipment volume using three years of historical monthly totals. The operations manager notes that volume has grown steadily and that December is always the busiest month. The data science consultant recommends a classical time series method that explicitly separates the long-term upward movement from the repeating yearly pattern. Which technique BEST fits this requirement?

A.k-means clustering on the monthly shipment totals
B.SARIMA (Seasonal AutoRegressive Integrated Moving Average)
C.A convolutional neural network trained on the monthly totals as a 1D signal
D.Logistic regression using the month number as the predictor
AnswerB

SARIMA extends ARIMA with seasonal autoregressive, differencing, and moving average terms, so it models trend and a repeating seasonal cycle simultaneously. With monthly data showing an upward trend and a strong December peak, the seasonal period of twelve captures the yearly pattern while the non-seasonal terms handle the growth. This makes SARIMA the direct match for the manager's stated need to separate trend from seasonality.

Why this answer

The scenario calls for decomposing a monthly series into a long-term trend and a twelve-month seasonal cycle, then projecting it forward. SARIMA is built precisely for that structure, using seasonal differencing and seasonal AR/MA terms alongside non-seasonal ones. Clustering has no forecasting capability, a CNN is impractical on thirty-six points and not interpretable as trend plus season, and logistic regression targets a categorical outcome rather than continuous volume.

Exam trap

The trap here is reaching for a flexible machine learning model by default, when classical time series methods like SARIMA are the correct tool for short series with clear trend and seasonality.

543
MCQhard

A team is training a deep learning model for image classification. The training loss decreases steadily but the validation loss plateaus after 20 epochs and then starts to increase. Which action is MOST likely to improve generalization?

A.Add more convolutional layers
B.Increase the learning rate
C.Implement early stopping
D.Reduce the batch size
AnswerC

Early stopping halts training at the epoch where validation loss begins rising, restoring the weights from the best validation checkpoint. This directly counters the stem's overfitting pattern, improving generalisation by preventing further memorisation of training data.

Why this answer

Early stopping halts training when validation loss stops improving, preventing overfitting. Increasing learning rate would worsen divergence; adding more layers increases capacity and overfitting; reducing batch size may help optimization but not directly address overfitting.

544
MCQeasy

A software company wants to add a feature that automatically transcribes customer support phone calls into text for analysis. Which type of AI technology is best suited for this task?

A.Computer vision
B.Natural language processing
C.Reinforcement learning
D.Automatic speech recognition
AnswerD

Automatic speech recognition (ASR) is specifically designed to convert spoken language into written text. It processes audio waveforms and outputs transcriptions, making it the ideal technology for transcribing customer support calls. Modern ASR systems handle various accents and background noise, and can be integrated with NLP for further analysis.

Why this answer

Automatic speech recognition (ASR) is the AI technology that converts spoken language into text. It is the foundational component for transcribing phone calls, enabling subsequent analysis. Computer vision handles images, NLP processes text after transcription, and reinforcement learning is for sequential decision-making, so none of those directly perform the speech-to-text conversion required.

Exam trap

The trap here is confusing natural language processing with speech recognition, assuming that any language-related task falls under NLP, when in fact audio-to-text conversion is a distinct domain.

545
MCQmedium

A retail company runs a demand-forecasting model in production. Over three weeks, the average order value of incoming transactions has risen by 40 percent because of a promotional campaign, and forecast error has grown steadily. The model was trained on twelve months of historical data with no promotion periods. Which action should the operations team take FIRST to restore forecast reliability?

A.Increase the inference batch size so the model processes more transactions per call.
B.Add a caching layer in front of the inference endpoint to reduce repeated calls.
C.Instrument input-feature drift monitoring and retrain or recalibrate the model with recent promotion-period data.
D.Lower the model's confidence threshold so more forecasts are emitted per period.
AnswerC

The promotional campaign shifted the distribution of input features away from the training distribution, which is data drift. Detecting it with feature-distribution monitoring and then retraining or recalibrating on data that includes promotion periods restores the mapping the model learned. This addresses the root cause rather than a symptom such as latency or throughput.

Why this answer

The rising average order value from the promotion changed the distribution of input features relative to the training set, which is classic data drift and explains the growing forecast error. Monitoring feature distributions detects the shift, and retraining or recalibrating with promotion-period data realigns the model with current conditions. Throughput, caching, and threshold changes do not alter the underlying statistical mismatch.

Exam trap

The trap here is assuming that a production accuracy problem must be fixed by tuning serving infrastructure or thresholds, when the evidence points to input data drift that only retraining or recalibration can correct.

546
Multi-Selectmedium

A media company exposes a text-to-image generation API built on a diffusion model. Users submit prompts and receive generated images. The security team wants to reduce the risk that the API can be abused to produce prohibited content such as realistic depictions of public figures in compromising situations. (Choose two.)

Select 2 answers
A.Deploy a prompt classifier that blocks or rewrites requests matching known prohibited-content patterns before generation.
B.Store all generated images indefinitely in an unencrypted bucket for later review by the security team.
C.Rate-limit each API key to a fixed number of image generations per minute.
D.Apply output filtering that scans generated images with a safety classifier and withholds or blurs those flagged as prohibited.
E.Enable TLS 1.3 for all API connections and require client certificates for authentication.
AnswersA, D

A prompt classifier inspects incoming text and can reject or sanitize requests that target prohibited subjects before the diffusion model ever runs. Blocking at this stage prevents the model from producing the disallowed image and reduces compute waste. It is a standard input-side guardrail for generative APIs and directly addresses abuse of the prompt channel.

Why this answer

Reducing abuse of a generative API requires controls that act on the content itself. A prompt classifier blocks or rewrites harmful requests before generation, and an output safety classifier catches disallowed images that still emerge. Together they provide input and output guardrails, while transport security, rate limiting, and insecure retention address different concerns and leave the content-abuse risk largely unmitigated.

Exam trap

The trap here is treating transport security or rate limiting as content moderation, when only prompt and output classifiers evaluate whether the material itself is prohibited.

547
MCQmedium

A data scientist is training a resume screening model to rank job applicants. The training data includes historical hiring decisions from the past 10 years. The company wants to avoid unfair bias against underrepresented groups. Which type of bias is most likely present in the training data?

A.Algorithmic bias
B.Selection bias
C.Confirmation bias
D.Historical bias
AnswerD

Historical bias arises because the model learns from a decade of past hiring decisions that already reflect prior discriminatory outcomes, so the training labels themselves encode the unfair patterns the company now wants to avoid.

Why this answer

Historical bias occurs when training data reflects past human decisions or societal patterns that were themselves biased, causing the model to learn and perpetuate those patterns. A resume screening model trained on 10 years of hiring decisions will inherit whatever demographic skews existed in those decisions.

Exam trap

AI0-001 often tests the distinction between historical bias (bias baked into the data by past human decisions) and algorithmic bias (bias from model design) — candidates frequently pick 'algorithmic bias' whenever a model produces unfair outcomes, regardless of root cause.

How to eliminate wrong answers

Option A is wrong because algorithmic bias refers to bias introduced by the algorithm's design, optimization objective, or feature engineering, not by the historical data itself. Option B is wrong because selection bias refers to non-random sampling of the data (e.g., only including applicants from certain sources), which is a data-collection issue rather than the inherited-decision issue described. Option C is wrong because confirmation bias is a human cognitive bias where people favor information confirming preexisting beliefs, not a property of training data.

548
Multi-Selectmedium

Which TWO actions should be taken to ensure an AI model complies with GDPR requirements when processing personal data?

Select 2 answers
A.Limit data collection to only what is necessary for the model
B.Provide a full explanation of model predictions
C.Store all user data for a minimum of 10 years
D.Anonymize all personal data before use
E.Implement user data deletion upon request
AnswersA, E

Data minimisation is a core GDPR principle, so restricting collection to what the model genuinely needs satisfies the lawfulness and storage-limitation requirements. This directly limits the volume of personal data processed, reducing exposure and helping demonstrate accountability under the regulation.

Why this answer

Option A is correct because GDPR's data minimization principle (Article 5(1)(c)) requires that personal data be adequate, relevant, and limited to what is necessary for the purposes for which it is processed, so restricting collection to only what the model needs directly supports compliance. Option E is correct because GDPR grants data subjects the right to erasure (Article 17, the 'right to be forgotten'), so implementing user data deletion upon request is a mandatory capability for any system processing personal data. Option B is not required by GDPR, which focuses on transparency about processing purposes and logic rather than mandating a full explanation of every model prediction (that concern relates more to interpretability and AI-specific regulation).

Option C is wrong because GDPR's storage limitation principle requires data to be kept no longer than necessary, so a 10-year minimum retention period would violate the regulation. Option D is not strictly required because GDPR permits processing of personal data with a lawful basis; anonymization is one way to reduce risk but is not a universal prerequisite, and true anonymization is often impractical for model training.

Exam trap

CompTIA often tests the misconception that anonymization is always required before any AI processing of personal data, but GDPR allows processing under lawful bases without anonymization, making Option D a tempting but incorrect choice.

549
MCQeasy

A developer wants to secure an AI API service. Which practice is MOST effective for preventing unauthorized access to the model?

A.Using a larger context window
B.Enforcing least-privilege API access with proper key management
C.Enabling response logging
D.Implementing rate limiting
AnswerB

Least-privilege API access scopes each key to only the operations its consumer needs, while proper key management enables rotation and revocation. Together these limit the blast radius of a leaked credential, directly preventing unauthorised access to the model.

Why this answer

Enforcing least-privilege API access with proper key management is the most effective practice because it ensures that each API key or token has only the minimum permissions necessary for its intended function, reducing the attack surface. Proper key management includes rotating keys, using scoped access tokens (e.g., OAuth 2.0 scopes), and storing keys securely (e.g., using a secrets manager like AWS Secrets Manager or HashiCorp Vault). This directly prevents unauthorized access by limiting what a compromised or misused key can do, unlike other options that address secondary concerns.

Exam trap

The AI0-001 exam often tests the distinction between preventive and detective controls, and the trap here is that candidates confuse rate limiting (a throttling mechanism) with access control, thinking it prevents unauthorized access when it only limits the frequency of requests.

How to eliminate wrong answers

Option A is wrong because using a larger context window increases the amount of input the model can process but does nothing to authenticate or authorize API requests; it is a model configuration parameter, not a security control. Option C is wrong because enabling response logging aids in auditing and detecting breaches after they occur, but it does not prevent unauthorized access in real time; it is a detective control, not a preventive one. Option D is wrong because implementing rate limiting mitigates denial-of-service attacks and abuse by throttling request volume, but it does not verify the identity or permissions of the requester; an attacker with a valid key could still access the model within rate limits.

550
MCQeasy

A startup wants to add an AI-powered virtual assistant to their mobile app. They have limited in-house AI expertise and need a solution that can be integrated quickly with minimal infrastructure management. Which deployment pattern is MOST suitable?

A.Implement an asynchronous processing queue for all user requests
B.Train and deploy a custom model on an on-premises server
C.Deploy the model on edge devices for offline inference
D.Use a cloud-based AI microservice (e.g., Amazon Lex, Azure Bot Service) with a pre-built model
AnswerD

A cloud-based AI microservice with a pre-built model removes the need to train, host or scale models, directly satisfying the limited in-house expertise and minimal infrastructure management constraints. Integration is largely API configuration, enabling rapid deployment within the mobile app.

Why this answer

Using a cloud-based AI microservice with a pre-built model is most suitable because it requires minimal AI expertise and infrastructure management. The startup can integrate the service via APIs quickly, leveraging the provider's pre-trained models and scalability. This aligns with their need for rapid integration and limited in-house AI skills.

Exam trap

The trap is overcomplicating the solution by assuming custom model training is necessary; the exam often tests the ability to choose managed services when expertise and time are limited.

How to eliminate wrong answers

Option A is wrong because an asynchronous processing queue is an architectural pattern for decoupling components, not a deployment pattern for AI assistants, and it does not address the need for pre-built AI capabilities. Option B is wrong because training and deploying a custom model on-premises requires significant AI expertise and infrastructure management, contradicting the startup's constraints. Option C is wrong because deploying on edge devices for offline inference is complex, requires model optimization, and may not be necessary for a mobile app that can use cloud services.

551
MCQmedium

A team trains a model to predict whether loan applicants will default. On the holdout set the model achieves 0.86 AUC, but when audited, applicants over 60 receive systematically higher risk scores than equally qualified younger applicants. The team must reduce this disparity while preserving predictive performance. Which action should they take first?

A.Remove the age column from the training data and retrain the model.
B.Increase model complexity so it can learn finer distinctions between applicants.
C.Lower the overall decision threshold to approve more applicants.
D.Measure disparity across protected groups and inspect whether age-correlated features drive the scores.
AnswerD

Before mitigating bias, the team must quantify it and locate its source. Computing group-level metrics such as demographic parity or equal opportunity, and examining feature importance for age-correlated variables, reveals whether the disparity comes from a proxy feature, label bias, or sampling. This diagnostic step is the prerequisite for any effective, targeted fairness intervention.

Why this answer

Bias mitigation should begin with measurement, because the source of disparity determines the right remedy. Group-level fairness metrics and feature attribution reveal whether age itself, a correlated proxy, or biased labels drive the scores. Only after diagnosing the cause can the team choose an appropriate intervention, such as reweighting, proxy removal, or post-processing.

Exam trap

The trap here is assuming that deleting the protected attribute eliminates bias, when correlated proxy features often preserve the disparity and can even mask it from casual inspection.

552
MCQmedium

A data science team uses Git for version control of model code and DVC for data versioning. They want to implement a model registry to track trained models, their hyperparameters, and performance metrics. Which tool is specifically designed for this purpose and integrates with the existing workflow?

A.Apache Airflow
B.Docker
C.MLflow Model Registry
D.Kubernetes
AnswerC

MLflow Model Registry is purpose-built to track trained models, their hyperparameters and performance metrics, and it integrates directly with Git and DVC workflows. It satisfies the stem's requirement for a dedicated registry rather than a general-purpose artefact store.

Why this answer

MLflow Model Registry is specifically designed for managing model versions, tracking metadata, and integrating with Git and DVC. Apache Airflow is for workflow orchestration, not model registry. Kubernetes is for container orchestration.

Docker is for containerization.

553
MCQeasy

An organisation is developing an AI policy. According to the NIST AI RMF, which function involves establishing policies and procedures to ensure the organisation governs AI responsibly?

A.Manage
B.Measure
C.Govern
D.Map
AnswerC

The Govern function establishes the policies, procedures, roles and accountability structures that ensure AI is developed and used responsibly across the organisation. It satisfies the stem's requirement for setting up governance, whereas Map, Measure and Manage handle risk identification, assessment and treatment.

Why this answer

The Govern function of the NIST AI RMF establishes policies, procedures, and organizational structures to ensure AI is developed and used responsibly. It is the foundational function that cultivates a risk-management culture and assigns accountability across the AI lifecycle.

Exam trap

AI0-001 often tests whether candidates can distinguish the four RMF functions by their verbs — Govern = policies/accountability, Map = context, Measure = analysis, Manage = treatment — and candidates frequently swap Govern and Manage.

How to eliminate wrong answers

Option A is wrong because Manage refers to allocating resources and implementing treatments to address risks identified by the other functions, not to establishing policies. Option B is wrong because Measure involves analyzing, assessing, and tracking AI risks using quantitative and qualitative methods. Option D is wrong because Map is about contextualizing and framing AI risks by understanding the system, its purpose, and its stakeholders.

554
MCQeasy

Which embedding type is MOST suitable for capturing semantic meaning of text in a RAG pipeline?

A.Bag-of-words vectors
B.Dense embeddings from a pre-trained transformer model
C.TF-IDF vectors
D.One-hot encoding
AnswerB

Pre-trained transformer embeddings encode contextual semantics into fixed dense vectors, so semantically similar passages land close together in vector space. This supports accurate similarity search and retrieval in a RAG pipeline, unlike sparse lexical vectors that capture term overlap rather than meaning.

Why this answer

Dense embeddings from pre-trained transformer models (e.g., BERT, Sentence-BERT) capture semantic meaning by representing text as high-dimensional vectors where similar meanings are close in vector space. This is essential for RAG pipelines to retrieve relevant documents based on semantic similarity rather than lexical overlap. Dense embeddings are the standard for modern semantic search.

Exam trap

The trap is confusing traditional sparse representations (TF-IDF, bag-of-words) with dense embeddings; candidates may think TF-IDF captures semantics because it weights terms, but it does not capture contextual meaning.

How to eliminate wrong answers

Option A is wrong because bag-of-words vectors ignore word order and semantics, relying only on word frequency, which fails to capture meaning. Option C is wrong because TF-IDF vectors are sparse and based on term frequency-inverse document frequency, which also ignores semantic relationships and context. Option D is wrong because one-hot encoding represents words as isolated categorical variables with no semantic similarity, making it unsuitable for semantic search.

555
MCQeasy

A startup is building a conversational AI assistant that must understand and generate human-like text. The team has limited labeled data and a modest budget for compute. They want to leverage existing large language models rather than pretraining one. Which approach best meets their needs?

A.Use a traditional n-gram language model trained on the startup's text data to generate responses.
B.Use a pretrained large language model via an API or open-source checkpoint and apply prompt engineering or lightweight fine-tuning for the assistant's domain.
C.Train a transformer model from scratch on the startup's proprietary conversation logs to ensure full control over the architecture.
D.Implement a rule-based chatbot using regular expressions and decision trees to handle user intents.
AnswerB

Leveraging a pretrained LLM avoids the enormous cost of pretraining. Prompt engineering requires no parameter updates, and lightweight fine-tuning methods like LoRA or adapter tuning adapt the model to the domain with minimal compute and data. This matches the startup's constraints while still delivering human-like understanding and generation, making it the most practical and cost-effective path.

Why this answer

For a startup with limited data and compute, the most effective strategy is to build on a pretrained large language model rather than starting from scratch. Pretrained LLMs already possess broad language understanding and generation capabilities. Prompt engineering can steer behavior without training, and parameter-efficient fine-tuning methods like LoRA adapt the model to the domain with minimal resources.

This balances quality, cost, and speed to deployment.

Exam trap

The trap here is assuming that training from scratch or using classical NLP methods can match the language understanding of pretrained LLMs, when the startup's constraints make leveraging pretrained models the only viable path.

556
MCQeasy

A municipality uses an AI system to triage requests for public housing assistance. Community advocates ask how they can challenge a denial that they believe resulted from an erroneous data match. Which governance mechanism is MOST appropriate to provide affected individuals with a route to contest automated outcomes?

A.A quarterly fairness audit that measures disparate impact across demographic groups in the triage outcomes.
B.A model card that documents the triage model's training data, performance metrics, and known limitations.
C.An appeals process that lets applicants request human review and correction of the data and logic behind a denial.
D.A published algorithmic impact assessment summarizing the system's risks and mitigations.
AnswerC

A contestability mechanism gives affected individuals a concrete path to question an automated outcome, have a human reconsider it, and correct erroneous inputs. This directly addresses the advocates' concern by ensuring decisions are not final simply because a model produced them. It also creates feedback that can reveal systematic data-matching errors, making it the most appropriate governance control for this scenario.

Why this answer

Contestability requires an actionable process through which an affected person can dispute an automated decision, obtain human review, and have errors corrected. An appeals process with human reconsideration provides that route. Impact assessments, model cards, and fairness audits improve transparency and systemic oversight but do not give individuals a means to challenge their own outcomes.

Exam trap

The trap here is equating transparency artifacts such as model cards or impact assessments with contestability, when disclosure alone does not give an individual a remedy.

557
MCQmedium

A company has a TensorFlow model trained on-premises and wants to deploy it on AWS SageMaker for scalable inference. What is the BEST way to package the model for deployment?

A.Convert the model to ONNX and upload to SageMaker
B.Upload the .h5 file to S3 and create a SageMaker endpoint directly
C.Package the model in a Docker container with a TensorFlow serving script and push to Amazon ECR
D.Use SageMaker Studio to train the model again from scratch
AnswerC

SageMaker deploys models from container images in Amazon ECR, so packaging the TensorFlow model with a serving script inside a Docker image provides the inference stack SageMaker requires. Plain model artefacts or notebooks cannot be served directly.

Why this answer

SageMaker expects models in a container format; the inference container should include the model artifacts and the serving code, allowing SageMaker to host it on scalable endpoints.

558
MCQeasy

An AI security team is conducting a threat model for a new document summarization service. They want to identify threats related to spoofing of the AI's identity. Which STRIDE category should they consider?

A.Repudiation
B.Tampering
C.Information disclosure
D.Spoofing
AnswerD

Spoofing covers impersonating a legitimate entity, so threats where an attacker or component masquerades as the AI service's identity fall squarely here. Mapping this to the summarisation service's authentication and identity claims satisfies the team's goal of identifying AI identity spoofing threats.

Why this answer

Spoofing is the STRIDE category that covers threats related to impersonating the AI's identity, such as an attacker pretending to be the AI service to gain unauthorized access or deceive users. The question explicitly asks about spoofing of the AI's identity.

Exam trap

The trap is that candidates might confuse Spoofing with Repudiation or Tampering, especially if they focus on the word 'identity' and think of authentication vs. authorization; Spoofing is specifically about impersonation.

How to eliminate wrong answers

Option A is wrong because Repudiation involves denying having performed an action, not impersonation. Option B is wrong because Tampering involves unauthorized modification of data or code, not identity spoofing. Option C is wrong because Information Disclosure involves unauthorized access to information, not identity impersonation.

559
MCQmedium

A healthcare analytics company has trained an AI model to predict patient readmission risk using a dataset that includes ZIP code, race, and historical healthcare costs. Before deployment, the compliance team runs a fairness audit and finds that the model's predictions correlate strongly with race even though race was not used as a direct input feature. Which of the following BEST describes this phenomenon?

A.Overfitting, because the model memorized race-related patterns in the training set instead of generalizing.
B.Label leakage, because the target variable inadvertently encodes race information.
C.Proxy discrimination, where a seemingly neutral feature acts as a substitute for a protected attribute.
D.Data drift, because patient demographics changed after the model was trained.
AnswerC

ZIP code and historical cost are correlated with race due to historical segregation and unequal access to care, so the model learns race-associated patterns indirectly. This is proxy discrimination: a protected attribute is not an explicit input, yet its influence enters through correlated features. Detecting it requires disparate impact testing and feature-correlation analysis, not just removing the protected column.

Why this answer

The model never receives race directly, but ZIP code and historical cost encode race-related disparities from decades of unequal healthcare access and residential segregation. This is proxy discrimination, and it defeats naive fairness approaches that only drop the protected column. Effective mitigation requires measuring disparate impact and auditing correlated features, because removing the explicit attribute does not remove its statistical footprint.

Exam trap

The trap here is assuming that removing the protected attribute from the feature set makes a model fair, when correlated proxy variables can preserve the same bias.

560
Multi-Selecthard

A team is implementing a RAG system for a legal document Q&A. They need to chunk documents effectively. Which THREE chunking strategies should they consider to improve retrieval accuracy for legal texts that contain hierarchical sections (clauses, sub-clauses, definitions)?

Select 3 answers
A.Hierarchical chunking that indexes chunks at clause and sub-clause levels with parent relationships
B.Overlapping chunks with a 10% overlap between consecutive chunks
C.Fixed-size chunking with a 512-token window and no overlap
D.Chunking based on the document's table of contents and section hierarchy
E.Semantic chunking that splits at natural boundaries (e.g., section headings, paragraph breaks)
AnswersA, D, E

Indexing at clause and sub-clause levels with parent links preserves the legal hierarchy, so retrieval can return a precise sub-clause while still supplying its enclosing clause as context. This directly addresses the stem's hierarchical sections constraint.

Why this answer

Option A is correct because hierarchical chunking indexes chunks at clause and sub-clause levels while preserving parent relationships, which matches the legal document's nested structure and lets retrieval return a precise sub-clause while still offering the enclosing clause for context. Option D is correct because chunking based on the document's table of contents and section hierarchy aligns chunk boundaries with the document's own logical divisions, so each chunk corresponds to a meaningful legal section rather than an arbitrary token span. Option E is correct because semantic chunking splits at natural boundaries such as section headings and paragraph breaks, keeping definitions and clauses intact and avoiding mid-sentence or mid-clause cuts that would degrade embedding quality and retrieval accuracy.

Option B is not among the correct answers because a fixed 10% overlap is a generic heuristic that does not respect legal hierarchy and can duplicate or fragment clause text. Option C is not among the correct answers because fixed-size 512-token windows with no overlap ignore section structure and will frequently split clauses, sub-clauses, and definitions across chunk boundaries.

561
MCQmedium

A data engineer is designing a data pipeline for a machine learning model that predicts equipment failure in a manufacturing plant. The pipeline ingests sensor data every second, and the model must be retrained daily with the latest data. The team needs to ensure that the training data is representative of the current operating conditions and that the model does not become stale. Which strategy is most appropriate for maintaining the training dataset?

A.Implement a sliding window approach where the training dataset includes only the most recent 30 days of sensor data.
B.Randomly sample 10% of all historical data for each daily retraining to reduce computational load.
C.Use the entire historical dataset from the plant's inception to train the model, ensuring maximum data volume.
D.Manually curate a fixed dataset of known failure events and use it for all retrainings.
AnswerA

A sliding window ensures that the training data reflects the most recent operating conditions, which is crucial when equipment behavior or environmental factors change over time. By using only the last 30 days, the model adapts to concept drift and remains relevant. This approach balances the need for sufficient data with the need for recency, and it is a common practice in time-series forecasting and predictive maintenance.

Why this answer

A sliding window of recent data is most appropriate because it keeps the training set aligned with current operating conditions, mitigating concept drift. This ensures the model learns from the most relevant examples, which is essential for predictive maintenance where equipment behavior evolves over time.

Exam trap

The trap here is assuming that more data is always better, but in dynamic environments, outdated data can actually harm model performance by introducing irrelevant patterns.

562
MCQhard

During inference, a model served via a REST API occasionally returns high latency due to cold starts. The team uses a containerized service on Kubernetes with horizontal pod autoscaling. Which solution minimizes cold start impact while controlling cost?

A.Configure the autoscaler based on request count with a shorter cooldown period
B.Increase CPU and memory requests for the inference container
C.Switch to vertical pod autoscaling
D.Use a sidecar container that pre-warms the model and set a minimum replica count
AnswerD

Pre-warming ensures the model is loaded; minimum replicas keep pods ready, reducing cold starts.

Why this answer

A sidecar warm-up agent and a minimum replica count keep pods ready. Increasing resources may not fix cold starts; autoscaling based on request count may lag; vertical scaling helps but not directly.

563
MCQeasy

A company wants to deploy a machine learning model that requires continuous learning as new data arrives. The model must be able to adapt to changing patterns without retraining from scratch. Which approach should be used?

A.Transfer learning
B.Online learning
C.Batch learning
D.Unsupervised learning
AnswerB

Online learning updates model parameters incrementally as each new data point arrives, rather than retraining on the full dataset. This satisfies the stem's constraint of continuous adaptation to changing patterns without retraining from scratch, unlike batch learning.

Why this answer

Online learning (also called incremental learning) updates the model incrementally as each new data point arrives, without requiring full retraining. This makes it ideal for scenarios where data arrives continuously and patterns shift over time, as the model can adapt its parameters on the fly.

Exam trap

CompTIA often tests the distinction between training paradigms (online vs. batch) and other ML concepts like transfer learning or unsupervised learning, so candidates may confuse 'continuous learning' with 'transfer learning' or incorrectly assume that any learning method can handle streaming data.

How to eliminate wrong answers

Option A is wrong because transfer learning reuses a pre-trained model on a new but related task, but it does not inherently support continuous adaptation to streaming data—it typically requires a separate fine-tuning phase. Option C is wrong because batch learning trains the model on the entire dataset at once and requires retraining from scratch when new data arrives, making it unsuitable for continuous learning. Option D is wrong because unsupervised learning is a paradigm for finding patterns in unlabeled data, not a deployment strategy for handling streaming data or model updates.

564
MCQhard

A research lab trains a language model using DP-SGD. What primary privacy risk does this technique mitigate?

A.Data poisoning attacks
B.Membership inference attacks
C.Adversarial patch attacks
D.Model inversion attacks
AnswerB

DP-SGD adds calibrated noise to per-example gradients during training, bounding any single record's influence on the model. This directly limits an adversary's ability to determine whether a specific individual's data was in the training set, satisfying the stem's requirement to mitigate membership inference attacks.

Why this answer

DP-SGD (Differentially Private Stochastic Gradient Descent) mitigates membership inference attacks by adding calibrated noise to gradients during training, which bounds the influence any single training example can have on the final model. This differential privacy guarantee makes it difficult for an adversary to determine whether a specific data point was included in the training set, directly addressing the core risk of membership inference.

Exam trap

The AI0-001 exam often tests the distinction between privacy risks (membership inference, model inversion) and security risks (poisoning, adversarial examples), and the trap here is that candidates confuse 'privacy risk' with 'security risk' and pick data poisoning or adversarial attacks instead of recognizing that DP-SGD is specifically designed for differential privacy against membership inference.

How to eliminate wrong answers

Option A is wrong because data poisoning attacks involve injecting malicious data to corrupt model behavior, which DP-SGD does not specifically prevent—it only limits per-example influence but does not detect or filter poisoned inputs. Option C is wrong because adversarial patch attacks target image classifiers by placing physical patches on objects to cause misclassification, which is a computer vision robustness issue unrelated to the privacy guarantees of DP-SGD. Option D is wrong because model inversion attacks aim to reconstruct training data features or attributes from the model, and while DP-SGD provides some defense, its primary and most direct mitigation is against membership inference, not full inversion which requires stronger assumptions and additional techniques.

565
MCQeasy

Which OWASP LLM Top 10 vulnerability involves an attacker manipulating the LLM through crafted inputs that override the system's intended instructions?

A.Sensitive information disclosure
B.Prompt injection
C.Supply chain vulnerabilities
D.Model denial of service
AnswerB

Prompt injection occurs when crafted input overrides the model's system instructions, hijacking its behaviour. It differs from insecure output handling or training-data poisoning, which target downstream execution or model weights rather than instruction hierarchy, directly matching the stem's override scenario.

Why this answer

Prompt injection (Option B) is the correct answer because it directly describes an attack where crafted inputs override the system's intended instructions, causing the LLM to execute unauthorized actions or reveal restricted information. This vulnerability exploits the LLM's inability to distinguish between user-supplied content and system-level directives, effectively hijacking the model's behavior.

Exam trap

CompTIA often tests candidates' ability to distinguish between the attack vector (prompt injection) and its potential outcomes (e.g., sensitive information disclosure), leading them to incorrectly select the consequence rather than the root vulnerability.

How to eliminate wrong answers

Option A is wrong because sensitive information disclosure is a consequence of other vulnerabilities (e.g., prompt injection or insecure output handling), not the mechanism of overriding instructions. Option C is wrong because supply chain vulnerabilities involve compromised third-party components (e.g., pre-trained models, libraries) rather than direct input manipulation. Option D is wrong because model denial of service focuses on exhausting computational resources (e.g., via excessive token generation or resource-intensive queries), not on subverting instruction adherence.

566
Multi-Selectmedium

A data engineering team is designing a data pipeline to process streaming sensor data and feed it into an ML model for anomaly detection. Which THREE components are essential for this pipeline?

Select 3 answers
A.Apache Airflow for scheduling recurring batch jobs
B.Amazon S3 as a data lake for storing raw sensor data
C.Snowflake as a real-time streaming destination
D.Apache Kafka for ingesting streaming sensor data
E.Apache Spark Structured Streaming for real-time processing
AnswersB, D, E

S3 is a scalable object store that can serve as a data lake for raw sensor data, accessible for both streaming and batch processing.

Why this answer

Amazon S3 is essential as a data lake for storing raw sensor data because it provides durable, scalable, and cost-effective object storage that can serve as a central repository for streaming data before and after processing. In a streaming pipeline, raw data must be persisted for reprocessing, historical analysis, and compliance, and S3's integration with Apache Spark and Kafka makes it a natural landing zone for sensor data.

Exam trap

CompTIA often tests the distinction between batch and streaming technologies, and the trap here is that candidates confuse Airflow's scheduling capability with real-time streaming orchestration, or assume Snowflake can act as a streaming sink when it is fundamentally a batch-oriented warehouse.

567
MCQmedium

You are an AI governance officer at a bank that uses a machine learning model to predict credit risk. The model was developed by an external vendor and uses a proprietary algorithm. The bank's compliance team has determined that the model must be explainable to meet regulatory requirements. However, the vendor claims the model is a 'black box' and cannot provide explanations. You need to ensure compliance while maintaining the model's performance. What is the best course of action?

A.Ignore the requirement as the model is proprietary
B.Ask the vendor to develop a custom explanation module
C.Replace the model with a simpler, interpretable model
D.Use a model-agnostic explanation technique like SHAP
AnswerD

SHAP is model-agnostic: it treats the vendor's proprietary model as a black box, approximating each feature's contribution to individual predictions from input-output behaviour alone. This delivers the per-decision explanations regulators require without retraining or replacing the model, preserving its performance.

Why this answer

D is correct because model-agnostic explanation techniques like SHAP (SHapley Additive exPlanations) can provide post-hoc interpretability for any black-box model without requiring access to its internal structure or proprietary algorithm. This allows the bank to meet regulatory explainability requirements while preserving the vendor's proprietary model and its predictive performance.

Exam trap

The trap here is that candidates may assume that a 'black box' model cannot be explained at all, leading them to choose replacement with a simpler model (Option C), when in fact model-agnostic techniques like SHAP or LIME can provide explanations without altering the model itself.

How to eliminate wrong answers

Option A is wrong because ignoring regulatory requirements is not a viable option for a financial institution; it would lead to non-compliance and potential penalties. Option B is wrong because asking the vendor to develop a custom explanation module would require the vendor to modify their proprietary algorithm, which they have stated is a 'black box' and cannot provide explanations, making this request impractical and likely impossible. Option C is wrong because replacing the model with a simpler, interpretable model would sacrifice the predictive performance that the current model provides, which may be critical for accurate credit risk assessment.

568
Multi-Selecteasy

Which TWO of the following are common activation functions used in neural networks? (Choose two.)

Select 2 answers
A.Gradient descent
B.LSTM
C.Dropout
D.ReLU
E.Sigmoid
AnswersD, E

ReLU (Rectified Linear Unit) is a standard activation function, outputting the input directly when positive and zero otherwise. Its piecewise-linear form satisfies the question's requirement for common neural network activations, alongside sigmoid and tanh, by enabling efficient gradient propagation during training.

Why this answer

ReLU (Rectified Linear Unit) is a widely used activation function that outputs the input directly if it is positive, and zero otherwise, introducing non-linearity while mitigating the vanishing gradient problem. Sigmoid is another common activation function that maps any real-valued input to a value between 0 and 1, making it useful for binary classification output layers. Both are fundamental building blocks in neural network architectures.

Exam trap

CompTIA often tests the distinction between activation functions and other neural network components like optimizers (gradient descent), architectures (LSTM), or regularization techniques (dropout), expecting candidates to recognize that only ReLU and Sigmoid directly compute a neuron's output from its input.

569
Multi-Selecthard

An AI operations team is monitoring a deployed image classification model. They notice a gradual increase in prediction confidence but a drop in accuracy. Which THREE actions should they take to diagnose the issue?

Select 3 answers
A.Analyze the model's calibration curve to see if confidence scores align with actual accuracy.
B.Increase the size of the training dataset by collecting more unlabeled data.
C.Compare the distribution of input features between training and recent production data.
D.Evaluate model performance on a held-out test set collected at deployment time.
E.Retrain the model immediately with the most recent data.
AnswersA, C, D

Rising confidence alongside falling accuracy signals overconfident misclassification, so plotting the calibration curve quantifies the gap between predicted probability and observed correctness. This reveals whether the model's probability outputs have drifted from true likelihoods, guiding recalibration or retraining decisions.

Why this answer

Option A is correct because a calibration curve (reliability diagram) directly compares predicted confidence against observed accuracy, revealing whether the model has become overconfident — exactly the symptom of rising confidence with falling accuracy. Option C is correct because comparing training versus production input feature distributions (e.g., via drift metrics like PSI or KL divergence) detects data drift or covariate shift that can degrade accuracy while inflating confidence. Option D is correct because evaluating on a held-out test set collected at deployment time establishes a stable baseline, distinguishing genuine model degradation from changes in the production data distribution.

Option B is not appropriate because collecting unlabeled data does not diagnose the cause and cannot be used for supervised evaluation without labels. Option E is not appropriate because immediate retraining without root-cause analysis risks masking the problem and may reinforce the drift or labeling issues causing the confidence-accuracy mismatch.

Exam trap

CompTIA often tests the distinction between diagnostic actions and corrective actions—candidates mistakenly jump to retraining (Option E) or data collection (Option B) instead of first analyzing calibration and data distribution (Options A, C, D) to identify the specific type of drift or miscalibration.

570
MCQmedium

An organization wants to implement an AI system to automatically categorize support tickets into predefined categories. They have a labeled dataset of 10,000 tickets. Which approach is MOST appropriate?

A.Use a rule-based system with keyword matching
B.Use a prompt-based LLM with few-shot examples
C.Fine-tune a pre-trained text classification model
D.Train a custom neural network from scratch
AnswerC

Fine-tuning a pre-trained text classification model leverages the 10,000 labelled tickets to adapt existing language representations to the organisation's specific categories, satisfying the stem's supervised categorisation requirement. This outperforms training from scratch on a small dataset and avoids the cost and latency of prompt-based large-model inference.

Why this answer

Fine-tuning a pre-trained text classification model is a standard and effective approach for supervised classification when labeled data is available.

571
MCQeasy

In the AI project lifecycle, which phase involves partitioning the dataset into training, validation, and test sets?

A.Data acquisition
B.Model selection
C.Data preparation
D.Problem definition
AnswerC

Data preparation covers cleaning, labelling and splitting the dataset into training, validation and test partitions before modelling begins. These subsets support fitting, hyperparameter tuning and unbiased final evaluation respectively, so the split belongs to this phase rather than data collection or deployment.

Why this answer

Data preparation includes splitting the data to evaluate model performance and prevent leakage.

572
MCQeasy

A company wants to use AI to automatically categorize customer support tickets into topics like 'billing', 'technical', 'account'. They have 10,000 labeled examples. Which algorithm is most suitable for this task?

A.DBSCAN
B.Apriori
C.Principal component analysis (PCA)
D.Logistic regression
AnswerD

Logistic regression is a supervised classifier that learns a decision boundary from the 10,000 labelled examples, outputting category probabilities. It satisfies the labelled multi-class text classification constraint, though multinomial extension is needed beyond binary billing versus non-billing decisions.

Why this answer

Logistic regression is a supervised learning algorithm that models the probability of a categorical outcome based on input features. With 10,000 labeled examples, it can efficiently learn decision boundaries to classify tickets into 'billing', 'technical', or 'account' by using a softmax (multinomial logistic regression) extension for multi-class classification.

Exam trap

CompTIA often tests the distinction between supervised and unsupervised learning, so the trap here is that candidates may confuse clustering (DBSCAN) or dimensionality reduction (PCA) with classification, overlooking that labeled data requires a supervised algorithm like logistic regression.

How to eliminate wrong answers

Option A is wrong because DBSCAN is an unsupervised clustering algorithm that groups data based on density, not classification; it cannot use labeled examples to predict predefined categories. Option B is wrong because Apriori is an association rule mining algorithm used for market basket analysis to find frequent itemsets, not for classifying text into topics. Option C is wrong because PCA is an unsupervised dimensionality reduction technique that transforms features to capture variance, but it does not perform classification or use labels to assign categories.

573
MCQmedium

A company is building an AI-powered document intelligence system to extract key fields from scanned invoices. The data contains 95% of invoices from one vendor and 5% from others. During model training, the F1 score is 0.95 on the overall test set, but the performance on the minority vendor invoices is very poor. What is the MOST likely cause?

A.The model is overfitting on the minority class
B.The dataset is imbalanced, and the model is biased toward the majority class
C.The data has a train/test leakage problem
D.The feature extraction is incorrect for the minority vendor invoices
AnswerB

With 95% of invoices from one vendor, the model optimises for the majority class, so minority-vendor patterns are underweighted during training. High overall F1 masks this, since majority-class performance dominates the metric. The poor minority-vendor results stem directly from this class imbalance, not from overfitting or feature scaling.

Why this answer

The F1 score of 0.95 on the overall test set is misleading because 95% of invoices come from one vendor, so the model can achieve high overall accuracy by performing well on the majority class while failing on the minority class. This is a classic class imbalance problem where the model is biased toward the majority class. The poor performance on minority vendor invoices confirms that the model has not learned to generalize across vendors.

Exam trap

AI0-001 often tests the misconception that a high overall F1 score guarantees good performance across all classes, but candidates must recognize that imbalanced datasets can mask poor minority class performance.

How to eliminate wrong answers

Option A is wrong because overfitting on the minority class would typically cause poor performance on the majority class, not the other way around; here the model performs well on the majority and poorly on the minority, indicating bias toward the majority. Option C is wrong because train/test leakage would inflate performance on both majority and minority classes, not selectively harm the minority class. Option D is wrong because incorrect feature extraction for the minority vendor would likely affect all classes if the features are shared, and the symptom is specifically poor performance on the minority class, which is explained by imbalance.

574
Multi-Selectmedium

Which THREE of the following are types of machine learning paradigms? (Choose three.)

Select 3 answers
A.Gradient boosting
B.Reinforcement learning
C.Unsupervised learning
D.Quantum computing
E.Supervised learning
AnswersB, C, E

Reinforcement learning involves an agent learning from rewards.

Why this answer

Reinforcement learning is a correct machine learning paradigm where an agent learns to make decisions by interacting with an environment, receiving rewards or penalties based on its actions. This trial-and-error approach is distinct from supervised and unsupervised learning, as it focuses on maximizing cumulative reward through exploration and exploitation.

Exam trap

CompTIA often tests candidates by listing specific algorithms (like gradient boosting) or adjacent technologies (like quantum computing) as distractors, hoping you confuse a technique or enabling technology with a fundamental learning paradigm.

575
MCQmedium

A company uses AI to generate marketing images. They want to ensure that the images are clearly identified as AI-generated to comply with transparency obligations. Which approach is most effective?

A.Add a disclaimer in the platform's terms of service
B.Include metadata in the image file indicating it is AI-generated
C.Embed a visible watermark stating 'AI-generated' in each image
D.Rely on deepfake detection algorithms to flag the images
AnswerC

A visible watermark embedded in each image directly satisfies transparency obligations, because viewers immediately recognise the content as AI-generated. This labelling persists within the image itself, unlike metadata, which can be stripped or overlooked during sharing.

Why this answer

A visible watermark embedded directly in each generated image is the most effective transparency measure because it travels with the image regardless of where it is copied, screenshotted, or re-shared, and it is immediately perceivable by any viewer. Metadata can be stripped, and terms of service are not seen by end viewers, so a visible watermark is the only option that guarantees the AI-generated nature is disclosed at the point of consumption.

Exam trap

The trap is assuming that metadata or terms of service satisfy 'transparency' — the exam expects you to recognize that only a visible, persistent marker guarantees the viewer actually sees the AI disclosure.

How to eliminate wrong answers

Option A is wrong because terms of service are legal documents that end viewers of an image never read — they do not provide transparency at the point of consumption. Option B is wrong because metadata (such as C2PA or EXIF tags) is easily stripped by social media platforms, screenshotting, or re-encoding, so it cannot guarantee disclosure. Option D is wrong because deepfake detection algorithms are reactive, imperfect, and used after the fact — they do not proactively label content as AI-generated and often produce false positives or negatives.

576
MCQmedium

An AI model is being developed for medical diagnosis from X-ray images. The dataset contains only frontal chest X-rays. The model achieves high accuracy on test set but fails on lateral views. What is the most likely cause?

A.Dataset bias
B.Underfitting
C.Label noise
D.Overfitting
AnswerA

Dataset bias arises because training data contains only frontal chest X-rays, so the model learns frontal-specific features and cannot generalise to lateral views. The constraint is the narrow training distribution, not model capacity or overfitting to the test set.

Why this answer

The model was trained exclusively on frontal chest X-rays, so it never learned features specific to lateral views. When tested on lateral views, the distribution shift causes poor performance, which is a classic case of dataset bias (sampling bias). The high accuracy on the test set is misleading because the test set also only contained frontal views, masking the model's inability to generalize to other X-ray orientations.

Exam trap

The AI0-001 exam often tests the distinction between overfitting and dataset bias by presenting a scenario where the model performs well on the test set (which shares the same bias) but fails on a different data distribution, leading candidates to mistakenly choose overfitting instead of recognizing the sampling bias.

How to eliminate wrong answers

Option B (Underfitting) is wrong because underfitting would cause poor performance on both the training and test sets, not just on a different distribution of data. Option C (Label noise) is wrong because label noise refers to incorrect ground-truth labels in the training data, which would degrade performance across all data types, not selectively on lateral views. Option D (Overfitting) is wrong because overfitting would manifest as high training accuracy but low test accuracy on the same distribution (e.g., frontal views), not as a failure on an entirely unseen data orientation.

577
Multi-Selectmedium

A machine learning engineer wants to track hyperparameter experiments and compare results across runs. Which TWO tools are best suited for this purpose? (Choose 2)

Select 2 answers
A.MLflow
B.Weights & Biases
C.Apache Airflow
D.Docker
E.Kubeflow
AnswersA, B

MLflow provides experiment tracking that logs parameters, metrics and artefacts per run, letting engineers compare hyperparameter configurations side by side. Its tracking server and UI directly satisfy the requirement to record and contrast results across multiple runs.

Why this answer

MLflow is correct because it provides a centralized tracking server and API to log hyperparameters, metrics, and artifacts for each run, enabling easy comparison across experiments. Weights & Biases is correct because it offers a cloud-hosted dashboard with real-time logging, hyperparameter sweeps, and collaborative comparison features, making it ideal for tracking and comparing runs.

Exam trap

CompTIA AI exams often test the distinction between infrastructure tools (orchestration, containerization) and purpose-built experiment tracking tools; the trap here is that candidates may confuse Kubeflow’s pipeline capabilities with dedicated experiment tracking, or assume Docker/Airflow can serve as tracking solutions because they are used in ML workflows.

578
MCQeasy

In the AI lifecycle, which phase involves splitting data into training, validation, and test sets?

A.Model training
B.Data preprocessing
C.Data collection
D.Model evaluation
AnswerB

Data preprocessing covers cleaning, transforming and partitioning the dataset, including the train/validation/test split. The split happens here because the model needs held-out data before training begins, ensuring validation tunes hyperparameters and the test set gives an unbiased final evaluation.

Why this answer

Data preprocessing is the phase where raw data is cleaned, transformed, and prepared for modeling. Splitting the dataset into training, validation, and test sets is a critical step during this phase to ensure unbiased evaluation and prevent data leakage. This split occurs before any model training begins, making it part of preprocessing rather than training or evaluation.

Exam trap

CompTIA often tests the misconception that data splitting belongs to model training or evaluation, when in fact it is a preprocessing step that must occur before any model sees the data.

How to eliminate wrong answers

Option A is wrong because model training is the phase where the algorithm learns patterns from the training data, not where the data is split; splitting must happen beforehand to avoid contaminating the evaluation. Option C is wrong because data collection is the initial gathering of raw data from sources, which occurs before any splitting or preprocessing. Option D is wrong because model evaluation uses the already-split test set to assess performance, but the split itself is established during data preprocessing.

579
MCQeasy

A model's training accuracy is 99% but validation accuracy drops to 60%. What is the most likely issue?

A.Data leakage
B.Overfitting
C.Multicollinearity
D.Underfitting
AnswerB

A large gap between 99% training accuracy and 60% validation accuracy means the model memorised training noise rather than generalising. Overfitting is the mechanism: high variance causes the model to fit patterns specific to the training set that do not hold on unseen validation data.

Why this answer

A training accuracy of 99% with a validation accuracy of only 60% is a classic symptom of overfitting. The model has memorized the training data, including noise and outliers, rather than learning generalizable patterns, causing it to perform poorly on unseen validation data.

Exam trap

CompTIA often tests the distinction between overfitting and data leakage by presenting a large accuracy gap, where candidates might mistakenly attribute the issue to data leakage instead of recognizing that leakage typically inflates both accuracies rather than creating a divergence.

How to eliminate wrong answers

Option A is wrong because data leakage typically causes both training and validation accuracy to be artificially high, not a large gap between them; it occurs when information from outside the training set inadvertently influences the model. Option C is wrong because multicollinearity refers to high correlation among input features in regression models, which affects coefficient stability and interpretability, not a drastic accuracy drop between training and validation sets. Option D is wrong because underfitting would result in low accuracy on both training and validation sets (e.g., both below 70%), not a high training accuracy with a low validation accuracy.

580
MCQmedium

A retail bank is deploying a customer-facing AI assistant that must never disclose internal policy text. The team has a system prompt with instructions, but red-team testing shows users can extract the policy by asking the model to 'repeat everything above this line.' Which implementation change most directly mitigates this prompt-injection extraction risk in production?

A.Reduce the context window size so there is less room for the model to repeat the system prompt.
B.Move the policy text from the system prompt into the fine-tuning dataset so the model learns it as weights.
C.Increase the model temperature so responses become less deterministic and harder to reverse engineer.
D.Add an input guardrail that detects and blocks prompt-extraction patterns before the request reaches the model.
AnswerD

An input guardrail inspects the user prompt for known injection and extraction patterns, such as 'repeat everything above,' and blocks or sanitizes the request before it reaches the model. This directly addresses the attack vector shown in red-team testing while keeping the system prompt private, and it can be tuned and logged without retraining the underlying model.

Why this answer

The extraction attack succeeds because untrusted user input is concatenated with trusted instructions and the model follows the most recent instruction. Placing a guardrail in front of the model detects and blocks known extraction phrasing, which is the most direct runtime mitigation. Temperature, fine-tuning, and context size do not intercept the malicious instruction before it is processed.

Exam trap

The trap here is assuming that hiding or shortening the system prompt prevents extraction, when the real fix is validating untrusted input before it reaches the model.

581
MCQmedium

A government agency is deploying an AI model to screen loan applications. The model uses features like income, credit score, employment history, and zip code. During fairness auditing, the model is found to deny a disproportionately high number of applicants from a particular demographic group, even when controlling for legitimate financial factors. The agency wants to mitigate this bias without significantly reducing overall accuracy. Which approach should the data scientist prioritize?

A.Adjust the decision threshold for the affected group
B.Remove the zip code feature from the model
C.Apply sample weighting to balance the demographic groups
D.Use adversarial debiasing during model training
AnswerD

Adversarial debiasing trains a predictor alongside an adversary that penalises demographic predictability, directly suppressing the zip-code proxy encoding group membership while retaining legitimate financial signal. This satisfies the stem's dual constraint: reducing disparate denial rates without materially sacrificing overall accuracy, unlike pre-processing fixes that discard predictive information.

Why this answer

Adversarial debiasing is the correct approach because it directly optimizes the model to remove sensitive information (e.g., demographic group membership) from its internal representations while preserving predictive accuracy. This technique trains a primary model to predict the target (loan approval) and an adversary to predict the protected attribute from the model's learned features, forcing the primary model to learn representations that are both accurate and unbiased. It addresses the root cause of bias—correlation between protected attributes and model predictions—without requiring post-hoc threshold adjustments or sacrificing overall performance.

Exam trap

The AI0-001 exam often tests the misconception that simply removing a sensitive feature (like zip code) or reweighting data is sufficient to eliminate bias, when in reality bias can be encoded through correlated proxies and requires algorithmic debiasing during training.

How to eliminate wrong answers

Option A is wrong because adjusting the decision threshold for the affected group is a post-hoc fairness intervention that can reduce accuracy for that group and may violate regulatory requirements for equal treatment; it does not address the underlying model bias. Option B is wrong because removing the zip code feature alone is insufficient—bias can still propagate through correlated features like income or employment history, and this approach may reduce model accuracy without guaranteeing fairness. Option C is wrong because sample weighting can help balance representation but does not prevent the model from learning biased correlations; it may also distort the training distribution and degrade accuracy on the majority group.

582
MCQeasy

A small e-commerce startup has only 800 labeled customer-support tickets and needs to classify new tickets into categories such as billing, shipping, and returns. The team has no budget for large-scale annotation and wants to leverage a model already trained on millions of general text documents. Which approach best fits this constraint?

A.Use a rule-based keyword matcher that assigns categories based on the presence of predefined terms.
B.Apply k-means clustering to the ticket text and label each cluster with the most frequent category.
C.Train a transformer from scratch on the 800 tickets using a high learning rate.
D.Fine-tune a pretrained language model on the 800 labeled tickets for the classification task.
AnswerD

Fine-tuning leverages representations learned from millions of general text documents, so the model already understands language structure and only needs task-specific adjustment. With 800 labeled examples, fine-tuning can achieve strong performance where training from scratch would fail. This approach fits the startup's limited annotation budget and directly addresses the multi-class ticket categorization need.

Why this answer

Fine-tuning a pretrained language model is ideal when labeled data is scarce, because the model already encodes general language knowledge and only needs adaptation to the ticket categories. This approach uses the 800 labels efficiently and outperforms training from scratch, rule-based matching, or unsupervised clustering.

Exam trap

The trap here is assuming that a small labeled dataset requires an unsupervised or rule-based workaround, when transfer learning from a pretrained model is specifically designed to succeed with limited task-specific labels.

583
MCQeasy

A data engineer needs to store training data in a format that supports columnar pruning during model training. Which storage format should they use?

A.Parquet
B.XML
C.JSON
D.CSV
AnswerA

Parquet stores data column-wise with per-column statistics and row-group metadata, letting readers skip irrelevant columns and row groups entirely. This columnar pruning reduces I/O during training, directly meeting the stem's requirement for efficient columnar access.

Why this answer

Parquet is the correct choice because it is a columnar storage format that enables column pruning, allowing the training process to read only the columns needed for model training rather than entire rows. This reduces I/O and speeds up data loading, which is critical for large-scale AI/ML workloads. Unlike row-oriented formats, Parquet stores data by columns, making it efficient for analytical queries and feature selection.

Exam trap

CompTIA often tests the misconception that JSON or CSV are acceptable for columnar pruning because they are common and human-readable, but the trap here is that only columnar formats like Parquet or ORC support efficient column-level access, while row-oriented formats require full record scans.

How to eliminate wrong answers

Option B (XML) is wrong because XML is a verbose, hierarchical text format that stores data row-wise and lacks columnar pruning capabilities, leading to high storage overhead and slow read performance for tabular data. Option C (JSON) is wrong because JSON is a row-oriented, self-describing format that requires parsing entire records even when only a subset of fields is needed, making it unsuitable for column pruning. Option D (CSV) is wrong because CSV is a flat, row-oriented text format that forces reading entire rows into memory, with no support for columnar storage or predicate pushdown, resulting in inefficient I/O for selective column access.

584
MCQmedium

A healthcare organization wants to use patient data to predict disease risk. They are concerned about bias in the model. Which step is most critical during the data preparation phase to mitigate bias?

A.Applying SMOTE to oversample minority classes
B.Using a more complex algorithm
C.Removing all demographic features
D.Ensuring the training data is representative of the target population
AnswerD

Bias originates in data whose composition misrepresents the population the model will serve. Ensuring the training set reflects the target population's demographics and disease prevalence prevents systematic underrepresentation, directly satisfying the data-preparation constraint before any modelling begins.

Why this answer

Ensuring the training data is representative of the target population is the most critical step during data preparation to mitigate bias because bias often originates from skewed or incomplete data that does not reflect the real-world distribution of patient demographics, conditions, and outcomes. Without a representative dataset, any subsequent preprocessing or modeling will propagate and potentially amplify existing disparities, leading to unfair or inaccurate predictions for underrepresented groups.

Exam trap

CompTIA often tests the misconception that bias can be fixed by technical tweaks like oversampling or removing sensitive attributes, when in fact the root cause is almost always unrepresentative training data that must be addressed at the collection or sampling stage.

How to eliminate wrong answers

Option A is wrong because SMOTE (Synthetic Minority Oversampling Technique) addresses class imbalance by generating synthetic samples for minority classes, but it does not correct for broader representational bias (e.g., missing demographic subgroups) and can introduce artifacts if the minority class itself is not representative of the target population. Option B is wrong because using a more complex algorithm does not mitigate bias; in fact, complex models can overfit to spurious correlations in biased data, making bias worse rather than reducing it. Option C is wrong because removing all demographic features can mask bias but does not eliminate it—protected attributes like race or age may be correlated with other features (e.g., zip code, medical history), leading to proxy discrimination, and this approach can also remove clinically relevant information needed for accurate risk prediction.

585
MCQmedium

A company built a speech-to-text model using a recurrent neural network (RNN). During deployment, the model performs poorly on accented speech. Which action would most effectively improve model robustness?

A.Collect a small sample of accented speech and fine-tune the model on that sample only.
B.Add dropout and reduce the number of RNN layers to prevent overfitting to the current data.
C.Augment the training dataset with various accented audio samples and retrain the model.
D.Replace the RNN with a convolutional neural network (CNN) for feature extraction.
AnswerC

Accented speech is underrepresented in the original training data, so the model never learnt those acoustic patterns. Augmenting with varied accented samples exposes the RNN to that distribution during retraining, directly improving robustness to accents.

Why this answer

Augmenting the training dataset with diverse accented audio samples directly addresses the root cause of poor performance—distribution shift between training and deployment data. Retraining the model on this enriched dataset allows the RNN to learn invariant features across accents, improving generalization without altering the model architecture or risking catastrophic forgetting from fine-tuning on a tiny sample.

Exam trap

CompTIA often tests the misconception that architectural changes (like switching to CNN or adding regularization) can fix data distribution mismatches, when the real solution is to address the missing data diversity through augmentation or retraining.

How to eliminate wrong answers

Option A is wrong because fine-tuning on a small sample of accented speech can cause catastrophic forgetting of the original training distribution and does not provide enough diversity to learn robust accent-invariant features. Option B is wrong because adding dropout and reducing layers addresses overfitting to the current data, but the core problem is underfitting to accented speech due to missing representative training examples, not overfitting. Option D is wrong because replacing the RNN with a CNN for feature extraction does not inherently solve the accent robustness issue; CNNs are effective for spatial patterns but less suited for sequential temporal dependencies in speech, and the fundamental problem remains the lack of accented training data.

586
MCQmedium

A security analyst is investigating a potential adversarial attack on a production image classifier. The attack involves tiny perturbations that are invisible to the human eye but cause the model to misclassify a stop sign as a speed limit sign. Which type of attack is this?

A.Data poisoning
B.Model inversion
C.Membership inference
D.Adversarial example
AnswerD

Adversarial examples are inputs deliberately perturbed by imperceptible amounts to force misclassification, exactly matching the stop-sign-to-speed-limit scenario. Unlike data poisoning, which corrupts training data, or model inversion, which extracts training information, this attack manipulates inference-time input pixels, exploiting the model's learned decision boundaries without altering the model itself.

Why this answer

This is an adversarial example attack, where imperceptible perturbations are added to the input (e.g., a stop sign) to cause the model to misclassify it (e.g., as a speed limit sign). The perturbations are crafted using gradient-based methods (like FGSM or PGD) to maximize the model's loss, exploiting its linearity in high-dimensional spaces. This differs from other attacks because it targets the inference phase, not the training data or model parameters.

Exam trap

Candidates often confuse adversarial examples with data poisoning because both involve modifying input data. However, adversarial examples are crafted during inference to cause misclassification, whereas data poisoning corrupts the training dataset to influence the model's learned behavior.

How to eliminate wrong answers

Option A is wrong because data poisoning involves corrupting the training dataset (e.g., injecting mislabeled samples) to compromise the model during training, not adding perturbations to a single input at inference time. Option B is wrong because model inversion attempts to reconstruct private training data from the model's outputs (e.g., generating a face from a facial recognition model), not to cause misclassification of a specific input. Option C is wrong because membership inference determines whether a particular data point was part of the training set by analyzing the model's confidence scores, not by altering an input to induce a misclassification.

587
MCQhard

A financial institution runs a credit-scoring model that must comply with internal governance requiring that every individual prediction be traceable to the input features that drove it, and that the explanation be produced at inference time for each applicant. The model is a complex gradient-boosted ensemble. Which approach best satisfies the requirement to generate a per-prediction explanation for each applicant?

A.Report feature importances from the trained ensemble
B.Apply SHAP values to each individual prediction
C.Increase the number of boosting rounds to improve model stability
D.Train a global surrogate decision tree on the ensemble's outputs
AnswerB

SHAP assigns each feature a contribution to a single prediction based on cooperative game theory, so every applicant receives an explanation showing which input values pushed the score up or down. It works with tree ensembles through efficient exact algorithms, and its additive guarantees make the per-prediction attribution auditable, matching the governance requirement precisely.

Why this answer

Per-prediction traceability requires a local explanation method that decomposes an individual score into feature contributions. SHAP provides exactly that by attributing the difference between a prediction and a baseline to each input feature, with efficient exact computation available for tree ensembles. Global summaries such as feature importance or surrogate trees describe aggregate behavior and cannot explain a single applicant's decision.

Exam trap

The trap here is treating global interpretability artifacts, such as feature importance or a surrogate tree, as if they explain individual predictions.

588
MCQmedium

Which AI governance framework is specifically designed by the U.S. National Institute of Standards and Technology (NIST) to help organizations manage AI risks?

A.ISO/IEC 27001
B.COBIT
C.NIST AI Risk Management Framework
D.GDPR
AnswerC

The NIST AI Risk Management Framework provides voluntary guidance structured around govern, map, measure and manage functions, specifically for managing AI risks. It is the NIST-published framework, unlike ISO/IEC 42001 or the EU AI Act, which originate elsewhere.

Why this answer

The NIST AI Risk Management Framework (AI RMF) is the specific governance framework developed by the U.S. National Institute of Standards and Technology to help organizations manage AI risks, including those related to trustworthiness, fairness, and robustness. It provides a structured approach for identifying, assessing, and mitigating risks throughout the AI lifecycle, aligning with NIST's role in setting standards for cybersecurity and risk management.

Exam trap

The AI0-001 exam often tests candidates by listing well-known frameworks or regulations (like ISO/IEC 27001 or GDPR) that are related to security or privacy but are not AI-specific, leading candidates to confuse general governance with the NIST AI RMF.

How to eliminate wrong answers

Option A is wrong because ISO/IEC 27001 is an international standard for information security management systems (ISMS), not an AI-specific risk management framework, and it focuses on general data security rather than AI risks. Option B is wrong because COBIT (Control Objectives for Information and Related Technologies) is a framework for IT governance and management, developed by ISACA, and does not address AI-specific risks or the NIST AI RMF. Option D is wrong because GDPR (General Data Protection Regulation) is a European Union regulation for data privacy and protection, not a risk management framework, and it is not designed by NIST.

589
Multi-Selecthard

A company is deploying a generative AI system that produces text content. To comply with emerging transparency obligations, which THREE measures should they implement?

Select 3 answers
A.Watermark AI-generated content
B.Disclose AI involvement to users
C.Encrypt all training data
D.Provide deepfake detection tools
E.Limit model access to internal employees only
AnswersA, B, D

Watermarking helps identify AI-generated content and is a transparency best practice.

Why this answer

Watermarking AI-generated content (Option A) is correct because it embeds an imperceptible, machine-detectable signal into the output, enabling provenance verification. This directly addresses transparency obligations by allowing downstream systems to identify synthetic text, which is a key requirement in emerging AI regulations such as the EU AI Act.

Exam trap

CompTIA often tests the distinction between security controls (encryption, access control) and governance/transparency measures, so candidates mistakenly select encryption or access restriction as fulfilling transparency obligations when they do not.

590
MCQmedium

A company wants to create an AI system that can identify objects in images. They have a large dataset of labeled images. Which type of neural network architecture is most suitable?

A.Transformer
B.Convolutional neural network (CNN)
C.Generative adversarial network (GAN)
D.Recurrent neural network (RNN)
AnswerB

CNNs apply convolutional filters that exploit spatial locality and translation invariance in pixel grids, learning hierarchical visual features. This suits labelled image data for object identification, where fully connected networks would ignore spatial structure and scale poorly.

Why this answer

Convolutional neural networks (CNNs) are specifically designed to process grid-like data such as images. They use convolutional layers to automatically learn spatial hierarchies of features (edges, textures, objects) from pixel data, making them the most suitable architecture for image classification tasks with labeled datasets.

Exam trap

CompTIA often tests the misconception that any 'neural network' can handle images equally, but the trap is that RNNs and Transformers are sequence-based and not optimized for spatial feature extraction, while GANs are generative, not discriminative.

How to eliminate wrong answers

Option A is wrong because Transformers are primarily designed for sequential data (e.g., text) using self-attention mechanisms; while they can be adapted for vision (Vision Transformers), they require large datasets and are not the standard choice for traditional image classification. Option C is wrong because Generative adversarial networks (GANs) are used for generating new data (e.g., synthetic images) rather than classifying or identifying objects in existing images. Option D is wrong because Recurrent neural networks (RNNs) are designed for sequential or time-series data (e.g., text, speech) and struggle with spatial relationships in images due to vanishing gradients and lack of translation invariance.

591
MCQmedium

A team is using a pre-trained BERT model for a sentiment analysis task on product reviews. They want to adapt it to their specific domain with limited labeled data. Which approach is MOST effective?

A.Use BERT as a feature extractor and train a logistic regression on top
B.Apply data augmentation to increase the dataset and then train from scratch
C.Train a new BERT model from scratch on the domain data
D.Fine-tune the pre-trained BERT model on the small labeled dataset
AnswerD

Fine-tuning updates BERT's pre-trained weights on the small labelled dataset, transferring general language representations to the sentiment domain. This suits limited labelled data far better than training from scratch, which would overfit the small sample.

Why this answer

Fine-tuning the pre-trained BERT model on the small labeled dataset is the most effective approach because BERT has already learned rich language representations from large-scale corpora. Fine-tuning updates all or some of the pre-trained weights on the target task, allowing the model to adapt to the domain with limited data. This transfer learning approach consistently outperforms feature extraction and training from scratch when labeled data is scarce.

Exam trap

AI0-001 often tests the misconception that feature extraction is equivalent to fine-tuning; candidates may choose the simpler feature-extraction approach, but the exam expects recognition that fine-tuning is more effective for domain adaptation with limited labeled data.

How to eliminate wrong answers

Option A is wrong because using BERT as a frozen feature extractor and training a logistic regression on top only leverages the pre-trained representations without adapting them to the domain, which typically yields lower performance than fine-tuning. Option B is wrong because training from scratch on augmented data still requires massive amounts of data and compute; augmentation cannot compensate for the lack of a large corpus, and training from scratch discards the benefits of pre-training. Option C is wrong because training a new BERT model from scratch on the domain data is infeasible with limited labeled data and would lead to severe overfitting and poor performance.

592
Multi-Selectmedium

A software company is developing an AI-powered code generation tool that suggests code snippets to developers. The company wants to align with the EU AI Act's transparency requirements. Which two actions should the company take? (Choose two.)

Select 2 answers
A.Provide documentation to developers about the AI system's capabilities and limitations
B.Publish the full training dataset used to train the code generation model
C.Ensure the AI system marks AI-generated code with a machine-readable watermark
D.Clearly disclose that the code suggestions are generated by an AI system and not by a human
E.Obtain explicit consent from every developer before they use the tool
AnswersA, D

The EU AI Act requires that providers of certain AI systems supply instructions for use, including information about the system's capabilities and limitations. For a code generation tool, documenting potential inaccuracies or biases helps developers use the tool responsibly and aligns with transparency obligations.

Why this answer

The EU AI Act's transparency obligations for AI systems that interact with users require clear disclosure that the user is interacting with an AI, and for providers to supply instructions for use that describe capabilities and limitations. For a code generation tool, informing developers that suggestions are AI-generated and documenting the tool's limitations fulfills these duties.

Exam trap

The trap here is conflating transparency with open-sourcing training data or obtaining consent, which are not required by the EU AI Act for this type of tool.

593
MCQmedium

A data scientist is training a binary classifier and observes that the training accuracy is 99% but the test accuracy is only 70%. Which of the following is the MOST likely cause?

A.The model is underfitting the training data
B.The learning rate is too high
C.The model is overfitting the training data
D.The test set contains data leakage from the training set
AnswerC

A 29-point gap between training and test accuracy is the classic signature of overfitting: the model has memorised training noise rather than learning generalisable patterns. The constraint in the stem is the large train-test performance divergence, which variance-reduction techniques such as regularisation or more data would address.

Why this answer

A large gap between training accuracy (99%) and test accuracy (70%) is the classic signature of overfitting: the model has memorized the training data, including its noise, and therefore fails to generalize to unseen examples. The high training score confirms the model has enough capacity to fit the training set, while the poor test score shows that capacity is being used for memorization rather than learning generalizable patterns.

Exam trap

AI0-001 often tests the train-vs-test accuracy gap pattern, and candidates confuse overfitting (high train, low test) with underfitting (low train, low test) or with data leakage (which inflates test scores).

How to eliminate wrong answers

Option A is wrong because underfitting produces low accuracy on BOTH training and test sets, not a 99% training score. Option B is wrong because an excessively high learning rate typically causes unstable or diverging loss curves, not a clean 99%/70% train-test split. Option D is wrong because data leakage from the test set into training would inflate test accuracy (making it suspiciously high), not depress it to 70%.

594
Multi-Selecthard

A team is deploying a deep learning model that uses a convolutional neural network (CNN) for image recognition. The model achieves high accuracy but is very slow to infer on edge devices. Which THREE optimization techniques should the team consider to speed up inference without significant accuracy loss? (Select three.)

Select 3 answers
A.Use larger convolutional filters (e.g., 7x7 instead of 3x3) to capture more context.
B.Use weight pruning to remove unnecessary connections in the network.
C.Implement knowledge distillation by training a smaller model to mimic the larger one.
D.Increase the number of convolutional layers to improve feature extraction.
E.Apply model quantization to reduce weight precision.
AnswersB, C, E

Pruning reduces computation and memory footprint.

Why this answer

Weight pruning removes redundant or less important connections (weights) from the neural network, reducing the number of computations required during inference. This directly speeds up inference on edge devices while typically causing only a minor drop in accuracy if done carefully, making it a standard optimization technique for deploying CNNs on resource-constrained hardware.

Exam trap

CompTIA often tests the misconception that increasing model capacity (larger filters or more layers) improves performance without considering the trade-off in inference speed, leading candidates to select options that actually worsen latency on edge devices.

595
Multi-Selectmedium

A financial services firm is deploying an LLM-based assistant that summarizes internal earnings reports. Compliance requires that the assistant's outputs be auditable and that sensitive financial figures not be sent to an external model provider. Which TWO implementation measures should the team adopt? (Choose two.)

Select 2 answers
A.Deploy the model in a private environment where inference occurs within the firm's network boundary.
B.Enable the model provider's zero-retention API option so prompts are not stored by the vendor.
C.Log each prompt, retrieved context, model version, and response with timestamps for later review.
D.Restrict the assistant to answering only yes-or-no questions about the earnings reports.
E.Increase the model's temperature to diversify summaries so no single output is treated as authoritative.
AnswersA, C

Running inference inside the firm's network prevents sensitive financial figures from leaving the organization, directly satisfying the data-residency requirement. It also gives the firm control over model versions and logging infrastructure, which supports auditability. External API calls would transmit the same sensitive content to a third party.

Why this answer

Auditability requires durable records of prompts, context, model versions, and responses, while data-residency requires that sensitive figures never leave the firm's boundary. Running inference in a private environment and logging every interaction together satisfy both requirements. The remaining options either still transmit data externally or do nothing for auditability.

Exam trap

The trap here is accepting a vendor-side zero-retention promise as equivalent to keeping data internal, when the figures are still transmitted to an external provider.

596
Multi-Selectmedium

A team is designing an AI agent that needs to interact with external APIs, search the web, and perform multi-step reasoning. Which TWO architectural components are essential for this agentic workflow? (Choose TWO.)

Select 2 answers
A.Fine-tuning the base model
B.ReAct pattern (Reasoning + Acting)
C.Tool use / function calling
D.Single-turn response generation
E.Static prompt with no iterations
AnswersB, C

The ReAct pattern interleaves reasoning traces with actions, letting the agent decide when to call an API or search the web and then feed results back into further reasoning. This directly satisfies the multi-step reasoning and external interaction requirements in the stem.

Why this answer

The ReAct pattern (Reasoning + Acting) is essential because it interleaves chain-of-thought reasoning steps with actions, allowing the agent to plan, invoke tools, observe results, and revise its plan across multiple steps—exactly what multi-step reasoning with external APIs and web search requires. Tool use / function calling is equally essential because it provides the mechanism for the model to invoke external APIs and web search functions with structured arguments and receive structured results back into the reasoning loop. Together, ReAct supplies the iterative reasoning-and-action control flow while function calling supplies the concrete interface to external systems.

Fine-tuning the base model is not required for this workflow, since tool use and reasoning patterns can be implemented via prompting and orchestration without retraining weights. Single-turn response generation is insufficient because the scenario demands multi-step iteration rather than one-shot answers. A static prompt with no iterations also fails, as the agent must dynamically observe tool outputs and loop through reasoning cycles.

Exam trap

AI0-001 often tests the misconception that fine-tuning or a larger model is the key to agentic behavior, when in fact the essential components are the reasoning-action loop and external tool integration.

597
Multi-Selectmedium

A logistics company is deploying an AI model that predicts delivery delays. The model will run on edge devices in trucks with intermittent connectivity. The team must ensure the deployment meets latency and reliability requirements. Which TWO implementation practices are MOST appropriate for this edge AI deployment? (Choose two.)

Select 2 answers
A.Disable model monitoring on the device to reduce CPU overhead and extend battery life.
B.Increase the model's parameter count so it can learn more complex delay patterns from historical data.
C.Quantize the model to a smaller numeric precision so it fits the device's memory and compute budget.
D.Implement local inference with store-and-forward caching so predictions continue offline and sync when connectivity returns.
E.Route every prediction request to the cloud so the model always uses the newest weights.
AnswersC, D

Quantization reduces model size and compute cost, which directly addresses the memory and latency constraints of edge hardware in trucks. It enables local inference without relying on a network round trip, supporting the intermittent-connectivity requirement. This is a standard optimization for constrained edge deployment and preserves acceptable accuracy when validated against the original model.

Why this answer

Edge deployment under intermittent connectivity requires the model to run locally and be small enough for the device. Quantization reduces size and compute, and local inference with store-and-forward caching keeps predictions available offline while queuing results for later sync. Cloud routing, larger models, and disabled monitoring all conflict with the latency, memory, and reliability constraints described.

Exam trap

The trap here is prioritizing model sophistication or cloud freshness when the binding constraints are on-device memory, latency, and offline operation.

598
MCQhard

A research lab is training a large language model on a cluster of GPUs. They notice that training throughput decreases significantly when scaling from 8 to 16 GPUs. The model uses data parallelism with synchronous updates. Which factor is most likely causing the decreased throughput?

A.The learning rate is too high for the larger effective batch size.
B.The model is not using mixed precision training.
C.Insufficient GPU memory causing out-of-memory errors.
D.Increased communication overhead for gradient synchronization across GPUs.
AnswerD

In synchronous data parallelism, gradients must be averaged across all GPUs after each backward pass. As the number of GPUs increases, the all-reduce communication cost grows, potentially becoming a bottleneck. This overhead can reduce throughput if the network bandwidth or latency is insufficient, especially when scaling from 8 to 16 GPUs.

Why this answer

Synchronous data parallelism requires all-reduce operations to synchronize gradients. As the number of GPUs increases, the communication volume and frequency grow, and if the interconnect bandwidth is limited, this becomes a bottleneck. This is a common scaling challenge in distributed training, leading to sublinear speedup or even decreased throughput.

Exam trap

The trap here is attributing throughput drops to model or hyperparameter issues rather than inter-GPU communication overhead.

599
MCQhard

A financial institution needs to integrate an AI-based credit scoring model into an existing mainframe system that processes transactions in COBOL. The model is deployed as a REST API. What is the best strategy to ensure minimal disruption and maintain data integrity?

A.Copy all transaction data to a cloud database for the model to access.
B.Use an API gateway with versioning and circuit breaker patterns.
C.Rewrite the mainframe system in Java to directly call the model.
D.Install a GPU on the mainframe to run the model natively.
AnswerB

An API gateway with versioning and circuit breakers isolates the COBOL mainframe from model changes, letting the REST API evolve without breaking existing calls. Circuit breakers halt failing requests, preventing cascading timeouts and preserving transaction data integrity during partial outages.

Why this answer

An API gateway with versioning and circuit breaker patterns allows the COBOL mainframe to call the REST API without modifying its core transaction logic. The gateway handles protocol translation, rate limiting, and failover, ensuring minimal disruption to the existing mainframe system while maintaining data integrity through transactional consistency and graceful degradation.

Exam trap

CompTIA often tests the misconception that legacy systems must be replaced or heavily modified to integrate with modern AI services, when in fact an API gateway provides a non-invasive integration layer that preserves the existing infrastructure.

How to eliminate wrong answers

Option A is wrong because copying all transaction data to a cloud database introduces latency, potential data inconsistency, and security risks, and does not address the integration between COBOL and the REST API. Option C is wrong because rewriting the mainframe system in Java is a massive, high-risk, and costly undertaking that would cause significant disruption and is unnecessary for integrating a REST API. Option D is wrong because installing a GPU on the mainframe does not enable native execution of the AI model, as mainframes lack the necessary software stack and the model is deployed as a REST API, not as a local executable.

600
Multi-Selecthard

A data science team is preparing a gradient boosting model to predict equipment failure from sensor data. They want to tune hyperparameters that primarily control model complexity and reduce overfitting. Which two hyperparameters should they focus on? (Choose two.)

Select 2 answers
A.Random seed for data shuffling
B.Output directory for saved model artifacts
C.Number of CPU threads used during training
D.Learning rate
E.Maximum tree depth
AnswersD, E

The learning rate scales each tree's contribution to the ensemble. A smaller value forces the model to build more trees and take smaller corrective steps, which typically reduces overfitting and improves generalization. Paired with an appropriate number of estimators, tuning the learning rate is a core complexity control in gradient boosting for noisy sensor data.

Why this answer

Learning rate and maximum tree depth are structural hyperparameters that govern how much each tree contributes and how complex each tree can become. Lowering them constrains the ensemble's capacity, which reduces variance and overfitting. The other choices affect runtime, reproducibility, or file storage rather than the bias-variance tradeoff.

Exam trap

The trap here is treating any configurable training setting as a hyperparameter for overfitting, when operational settings such as thread count, seed, and output paths do not change model complexity.

Page 7

Page 8 of 13

Page 9