Courseiva

CCNA Implementing AI Solutions Questions

57 of 132 questions · Page 2/2 · Implementing AI Solutions · Answers revealed

76
Multi-Selectmedium

A team is designing an AI agent that needs to interact with external APIs, search the web, and perform multi-step reasoning. Which TWO architectural components are essential for this agentic workflow? (Choose TWO.)

Select 2 answers
A.Fine-tuning the base model
B.ReAct pattern (Reasoning + Acting)
C.Tool use / function calling
D.Single-turn response generation
E.Static prompt with no iterations
AnswersB, C

The ReAct pattern interleaves reasoning traces with actions, letting the agent decide when to call an API or search the web and then feed results back into further reasoning. This directly satisfies the multi-step reasoning and external interaction requirements in the stem.

Why this answer

The ReAct pattern (Reasoning + Acting) is essential because it interleaves chain-of-thought reasoning steps with actions, allowing the agent to plan, invoke tools, observe results, and revise its plan across multiple steps—exactly what multi-step reasoning with external APIs and web search requires. Tool use / function calling is equally essential because it provides the mechanism for the model to invoke external APIs and web search functions with structured arguments and receive structured results back into the reasoning loop. Together, ReAct supplies the iterative reasoning-and-action control flow while function calling supplies the concrete interface to external systems.

Fine-tuning the base model is not required for this workflow, since tool use and reasoning patterns can be implemented via prompting and orchestration without retraining weights. Single-turn response generation is insufficient because the scenario demands multi-step iteration rather than one-shot answers. A static prompt with no iterations also fails, as the agent must dynamically observe tool outputs and loop through reasoning cycles.

Exam trap

AI0-001 often tests the misconception that fine-tuning or a larger model is the key to agentic behavior, when in fact the essential components are the reasoning-action loop and external tool integration.

77
Multi-Selectmedium

A logistics company is deploying an AI model that predicts delivery delays. The model will run on edge devices in trucks with intermittent connectivity. The team must ensure the deployment meets latency and reliability requirements. Which TWO implementation practices are MOST appropriate for this edge AI deployment? (Choose two.)

Select 2 answers
A.Disable model monitoring on the device to reduce CPU overhead and extend battery life.
B.Increase the model's parameter count so it can learn more complex delay patterns from historical data.
C.Quantize the model to a smaller numeric precision so it fits the device's memory and compute budget.
D.Implement local inference with store-and-forward caching so predictions continue offline and sync when connectivity returns.
E.Route every prediction request to the cloud so the model always uses the newest weights.
AnswersC, D

Quantization reduces model size and compute cost, which directly addresses the memory and latency constraints of edge hardware in trucks. It enables local inference without relying on a network round trip, supporting the intermittent-connectivity requirement. This is a standard optimization for constrained edge deployment and preserves acceptable accuracy when validated against the original model.

Why this answer

Edge deployment under intermittent connectivity requires the model to run locally and be small enough for the device. Quantization reduces size and compute, and local inference with store-and-forward caching keeps predictions available offline while queuing results for later sync. Cloud routing, larger models, and disabled monitoring all conflict with the latency, memory, and reliability constraints described.

Exam trap

The trap here is prioritizing model sophistication or cloud freshness when the binding constraints are on-device memory, latency, and offline operation.

78
MCQeasy

A hospital is deploying an AI triage assistant that suggests priority levels for emergency room patients. Clinicians will review every suggestion before acting. The compliance team requires that the system log who reviewed each suggestion, what the clinician decided, and whether they overrode the AI. Which implementation practice best satisfies this requirement?

A.Store clinician feedback in a separate quality-improvement database that is refreshed monthly.
B.Enable automatic acceptance of AI suggestions when the model's confidence score exceeds 0.95.
C.Record a human-in-the-loop audit trail that captures the AI suggestion, the clinician's decision, and any override reason.
D.Log only the model's input features and output priority so that predictions can be reproduced later.
AnswerC

A human-in-the-loop audit trail preserves the full decision context: what the model proposed, what the clinician chose, and why any divergence occurred. This supports accountability, post-deployment review, and regulatory inspection. It also enables monitoring of override rates as a signal of model drift or poor fit to clinical workflow.

Why this answer

Human oversight in high-stakes AI requires evidence that a qualified person reviewed each suggestion and retained authority to disagree. An audit trail that links the suggestion, the human decision, and the override rationale provides that evidence and supports continuous monitoring. Auto-acceptance, model-only logging, and delayed batch feedback all fail to document meaningful human involvement.

Exam trap

The trap here is treating model reproducibility logs as equivalent to human-in-the-loop accountability records.

79
Multi-Selectmedium

A team is designing a RAG system for a large collection of PDFs. They need to choose document chunking strategies. Which TWO strategies are considered best practices? (Choose two.)

Select 2 answers
A.Semantic chunking (e.g., sentence or paragraph boundaries)
B.Fixed-size chunking with no overlap
C.Hierarchical chunking (sections, subsections)
D.Single chunk per document
E.Random character-length chunks
AnswersA, C

Semantic chunking splits PDFs at sentence or paragraph boundaries, so each chunk carries one coherent idea. This satisfies the retrieval-quality constraint: embeddings represent complete propositions, avoiding the mid-sentence fragmentation that degrades similarity matching in a RAG pipeline.

Why this answer

Semantic chunking (A) is a best practice because splitting text at natural sentence or paragraph boundaries preserves coherent, self-contained units of meaning, which improves embedding quality and retrieval relevance in a RAG pipeline. Hierarchical chunking (C) is also a best practice because it captures the document's section and subsection structure, allowing retrieval at multiple granularities (e.g., retrieving a subsection but supplying its parent section as context) and better handling long, structured PDFs. Fixed-size chunking with no overlap (B) is not recommended here because it can cut sentences or ideas mid-thought and provides no overlap to preserve context across boundaries.

A single chunk per document (D) is unsuitable because large PDFs would exceed embedding model token limits and dilute semantic focus, hurting retrieval precision. Random character-length chunks (E) are arbitrary and break semantic and structural coherence, making retrieval unreliable.

Exam trap

AI0-001 often tests chunking best practices, and candidates mistakenly believe fixed-size or single-chunk approaches are simpler and therefore acceptable, missing that semantic and hierarchical strategies preserve meaning and structure.

80
MCQmedium

A team is considering whether to fine-tune a base LLM or use RAG for a question-answering system over a large, static corpus of scientific papers. The answer must be highly accurate and grounded in the papers. Which approach is BEST and why?

A.Fine-tuning because it adapts the model to the scientific domain
B.Fine-tuning because it is faster at inference time
C.RAG because it retrieves and grounds answers in the source documents
D.RAG because it does not require any labeled data
AnswerC

RAG retrieves relevant passages from the static corpus and conditions generation on them, grounding answers in the source papers. Fine-tuning bakes knowledge into weights, which risks hallucination and staleness. The requirement for accuracy grounded in the documents makes retrieval the fitting choice.

Why this answer

RAG is the best fit because it retrieves relevant passages from the scientific corpus at query time and injects them into the LLM's context, grounding the answer in the actual source documents. This directly satisfies the requirement for high accuracy and traceability to the papers, and it avoids the cost and staleness issues of fine-tuning on a large static corpus. Fine-tuning changes model weights but does not guarantee the model will cite or stay faithful to specific documents.

Exam trap

The trap is assuming fine-tuning is always better for domain specialization, when the question's emphasis on grounding in source documents points decisively to RAG.

How to eliminate wrong answers

Option A is wrong because fine-tuning adapts the model's style and domain vocabulary but does not provide retrieval grounding — the model can still hallucinate facts not present in its weights, and updating the corpus requires retraining. Option B is wrong because fine-tuning does not make inference faster; in fact, serving a fine-tuned model requires the same or greater compute, and the speed argument is irrelevant to the accuracy requirement. Option D is wrong because while RAG does reduce the need for labeled data, that is a secondary benefit — the primary reason to choose RAG here is grounding and citation, not the absence of labels.

81
MCQhard

A hospital is deploying a vision model that flags possible pneumonia on chest radiographs. Radiologists report that the model performs well overall but frequently flags images from a newly installed portable X-ray unit. The images are technically adequate. The team must diagnose the cause before changing the model. Which action should the team take FIRST?

A.Retrain the model on a larger dataset that includes images from the new portable unit.
B.Add a rule that discards any image whose DICOM metadata indicates the portable unit was used.
C.Compare the statistical distribution of pixel intensities, resolution, and metadata between images from the new unit and the original training set.
D.Lower the model's classification threshold so fewer images are flagged as positive.
AnswerC

Comparing input distributions identifies distribution shift, which is the most likely cause when a model degrades on images from a new acquisition device. This diagnostic step reveals whether preprocessing, normalization, or resolution differences explain the false positives before any retraining, and it is non-destructive and reversible.

Why this answer

When a model degrades on inputs from a new device, the first step is to characterize the input distribution and compare it with training data. That comparison can reveal distribution shift caused by differences in resolution, exposure, or preprocessing, which must be understood before retraining or adjusting thresholds. Only after the cause is identified can the team choose an appropriate remediation.

Exam trap

The trap here is jumping to retraining or threshold changes when the scenario asks for the FIRST diagnostic step, which should isolate whether the new unit's images differ from the training distribution.

82
MCQmedium

A team is evaluating an LLM-based chatbot that frequently hallucinates when answering questions about internal policies. Which testing approach would MOST effectively quantify this issue?

A.Evaluation frameworks for LLM output quality
B.Integration tests for API calls
C.Unit tests for the data pipeline
D.Regression testing of model accuracy over time
AnswerA

Evaluation frameworks score LLM outputs against ground-truth policy answers using metrics such as groundedness and faithfulness, producing a repeatable hallucination rate. This quantifies the issue, satisfying the stem's requirement for measurable frequency rather than anecdotal review.

Why this answer

Evaluation frameworks for LLM output quality, such as those using metrics like faithfulness, factuality, or ROUGE/BLEU scores, are specifically designed to detect and quantify hallucinations by comparing generated responses against a ground-truth knowledge base. This directly measures the rate at which the chatbot fabricates or misstates internal policy details, providing a quantitative baseline for improvement.

Exam trap

The AI0-001 exam often tests the distinction between functional testing (e.g., API integration, data pipeline) and output quality evaluation, leading candidates to mistakenly choose integration or unit tests when the real issue is semantic accuracy of generated content.

How to eliminate wrong answers

Option B is wrong because integration tests for API calls verify that the chatbot's endpoints and external service interactions work correctly, but they do not assess the semantic accuracy or factual consistency of the generated text. Option C is wrong because unit tests for the data pipeline validate data ingestion, transformation, and storage logic, not the output quality of the LLM's responses. Option D is wrong because regression testing of model accuracy over time typically measures performance on a static benchmark (e.g., classification accuracy) rather than quantifying open-ended hallucination rates in a conversational context.

83
MCQhard

A hospital's AI triage assistant was validated on data from its own emergency department. Before rolling it out to three affiliated hospitals with different patient demographics, imaging equipment, and documentation habits, the governance committee requires evidence that the model will not silently underperform at the new sites. Which activity BEST provides that evidence?

A.Fine-tune the model on a sample of records from the three new hospitals before any prospective evaluation.
B.Run an external validation using held-out data from each receiving hospital and compare subgroup performance against the original site.
C.Increase the model's confidence threshold at the new sites until the override rate matches the original hospital's rate.
D.Re-run the original internal test set and confirm that the AUC is unchanged from the validation report.
AnswerB

External validation on each target site's own held-out data directly measures whether performance generalizes across demographics, equipment, and documentation differences, and subgroup analysis exposes disparities that aggregate accuracy would hide. This is the accepted method for pre-deployment generalization evidence in clinical AI governance. The other approaches either do not test the new populations or cannot reveal site-specific failure.

Why this answer

Generalization to new sites can only be demonstrated with data the model has never seen from those sites. External validation on each hospital's held-out records, broken down by subgroup, reveals demographic, equipment, and workflow-related performance gaps before patients are exposed. Reusing the internal test set, tuning thresholds, or fine-tuning prematurely all fail to produce that evidence and would leave the committee without a defensible basis for approval.

Exam trap

The trap here is treating a strong internal validation AUC as proof of generalization, when internal test data shares the exact site characteristics that differ at the new hospitals.

84
MCQhard

An organization runs a customer-support LLM that calls internal tools to look up order status and issue refunds. Security testing reveals that a user can paste text into the chat that causes the model to invoke the refund tool with an attacker-controlled amount. The team wants to reduce this prompt-injection risk without removing tool functionality. Which control is MOST effective?

A.Enforce authorization and parameter validation in the tool backend so refund requests are validated against the authenticated user's entitlements and business limits.
B.Fine-tune the model on a curated dataset of injection attempts so it learns to recognize and refuse malicious prompts.
C.Increase the model's temperature to zero so that responses become deterministic and injection attempts produce consistent refusals.
D.Add a system prompt instruction telling the model to ignore any user instructions that attempt to change its tool-use policy.
AnswerA

Placing authorization and validation in the tool backend creates a deterministic control that the model cannot talk its way past, because the refund service independently checks the caller's identity, order ownership, and amount limits. This defense-in-depth approach assumes the LLM may be manipulated and ensures that even a successful injection cannot exceed the user's actual entitlements, which is the standard pattern for agentic AI.

Why this answer

Prompt injection cannot be fully solved at the model layer, so the durable control is to enforce authorization and business rules in the tool backend where the model cannot influence them. Validating the authenticated user's entitlements and refund limits means a manipulated model still cannot perform unauthorized actions. Prompt instructions, temperature changes, and fine-tuning all rely on model compliance and fail as security boundaries.

Exam trap

The trap here is assuming prompt-level defenses such as system instructions or fine-tuning can serve as a security boundary for tool calls, when enforceable controls belong in the backend.

85
MCQeasy

A retail company wants to use AI to personalize marketing emails. They have a large dataset of customer purchase history and demographics. The data science team plans to use a collaborative filtering approach. Which data is MOST critical for this approach?

A.Email open rates and click-through rates from previous campaigns.
B.Customer purchase history and product ratings.
C.Product descriptions and categories.
D.Customer demographic information such as age and gender.
AnswerB

Collaborative filtering relies on user-item interactions, such as purchases or ratings, to find similarities between users or items. Purchase history provides implicit feedback, while ratings provide explicit feedback. This data is essential to generate recommendations based on patterns of co-occurrence or similarity, making it the most critical for the approach.

Why this answer

Collaborative filtering algorithms, such as matrix factorization or nearest neighbors, require user-item interaction data to identify patterns. Purchase history and ratings are the quintessential interaction data, enabling the system to recommend products based on similar users' behavior or similar items' co-occurrence. Other data types are either for content-based filtering or auxiliary.

Exam trap

The trap here is confusing collaborative filtering with content-based filtering, which uses item features like descriptions.

86
MCQhard

A financial services firm is implementing an AI solution that scores loan applications. The model must be auditable, and regulators require the firm to explain why any individual application received a particular decision. The data science team trained a gradient-boosted tree model with high accuracy. Which approach best meets the explainability requirement for individual decisions?

A.Replace the gradient-boosted tree with a logistic regression model and report the model coefficients as the explanation for every decision.
B.Provide the raw input features and the final score to the regulator and let them interpret the decision themselves.
C.Report the model's global feature importance ranking from the training run as the explanation for each decision.
D.Use SHAP values computed for each individual application to attribute the model's output to its input features, and present those attributions as the explanation.
AnswerD

SHAP values provide a per-instance, additive attribution of the model output to each input feature, which is exactly what is needed to explain an individual decision. They work with gradient-boosted trees and other complex models, and they are grounded in cooperative game theory, giving a consistent and locally accurate explanation. This satisfies the auditability requirement without sacrificing model accuracy, and the attributions can be logged and reviewed.

Why this answer

Regulatory explainability for individual decisions requires a per-instance attribution method, not a global summary or a raw score. SHAP values attribute the model output for a single application to its input features, preserving the accuracy of the gradient-boosted tree while producing a defensible, decision-specific explanation. Global importance, model replacement, and raw disclosure all fail to explain the specific decision under review.

Exam trap

The trap here is confusing global feature importance with local, per-instance explanations, or assuming that a simpler model alone satisfies a decision-specific audit requirement.

87
MCQmedium

A recommendation system for an e-commerce site is producing stale suggestions that do not reflect recent user behavior. The system is updated offline every 24 hours. Which change would MOST directly address this issue?

A.Increase the number of features used in the model
B.Add more training data from the past year
C.Use a deeper neural network architecture
D.Implement online learning to update the model incrementally in real time
AnswerD

Online learning updates model parameters incrementally as each interaction arrives, so recommendations reflect recent behaviour within seconds rather than waiting for the 24-hour offline batch. This directly removes the staleness constraint described in the stem.

Why this answer

Implementing online learning to update the model incrementally in real time directly addresses the staleness issue by allowing the model to incorporate recent user behavior as it happens. Online learning updates model parameters continuously or at short intervals, so recommendations reflect the latest interactions. This is the most direct solution to the problem of a 24-hour offline update cycle.

Exam trap

AI0-001 often tests the misconception that more data or a more complex model will solve staleness, when the core issue is the update frequency, and the most direct fix is to reduce latency through online learning.

How to eliminate wrong answers

Option A (Increase the number of features used in the model) is wrong because adding features does not address the latency of updates; the model would still be stale between updates. Option B (Add more training data from the past year) is wrong because more historical data does not solve the staleness problem; it may even reinforce old patterns. Option C (Use a deeper neural network architecture) is wrong because a deeper model does not inherently reduce update latency; it may improve accuracy but not freshness.

88
MCQmedium

A developer is building an AI microservice that processes document intelligence requests asynchronously. Users upload PDFs, and the service extracts text and analyzes it with an LLM. The processing time per document can be up to 5 minutes. Which integration pattern is MOST appropriate?

A.Synchronous REST API call that waits for the LLM response
B.Async processing with a message queue and separate worker service
C.WebSocket connection for real-time streaming
D.Serverless function triggered by HTTP request
AnswerB

A message queue decouples upload from processing, letting a separate worker service handle documents that may take five minutes each without blocking the API or hitting request timeouts. This matches the stem's asynchronous, long-running processing constraint.

Why this answer

Async processing with a message queue and separate worker service is the correct pattern because document processing can take up to 5 minutes, which exceeds typical synchronous HTTP timeout limits (often 30-60 seconds). The API accepts the upload, enqueues a job, and returns immediately; a worker service consumes the queue, performs text extraction and LLM analysis, and stores results for later retrieval. This decouples request handling from long-running work and provides resilience against worker failures.

Exam trap

The trap is underestimating HTTP and serverless timeout limits — candidates pick synchronous or serverless options without accounting for the 5-minute processing time exceeding those constraints.

How to eliminate wrong answers

Option A is wrong because a synchronous REST call waiting 5 minutes will hit client, load balancer, and API gateway timeouts, and it ties up server resources for the entire duration. Option C is wrong because WebSockets are for bidirectional real-time streaming, not for asynchronous batch document processing — they add complexity without solving the timeout problem. Option D is wrong because a serverless function triggered by HTTP is still synchronous from the caller's perspective and is subject to execution time limits (e.g., AWS Lambda's 15-minute max, but API Gateway's 29-second timeout), making it unsuitable for 5-minute LLM calls without additional async patterns.

89
MCQeasy

A company is building a document intelligence system that extracts key fields from scanned invoices. They have a labeled dataset of 10,000 invoices but need to decide between a traditional OCR+rule-based pipeline and an AI-based model. Which use case characteristic STRONGLY favors the AI-based approach?

A.Invoice layouts vary significantly between different vendors and often change
B.The system must process invoices in real time with sub-second latency
C.The team has limited access to labeled training data
D.Invoices have a fixed, standardized layout across all vendors
AnswerA

Varying and frequently changing vendor layouts defeat fixed templates and hand-written extraction rules, which need constant rework. An AI model learns visual and textual patterns from the 10,000 labelled invoices, generalising to unseen layouts, so this characteristic strongly favours the AI-based approach.

Why this answer

An AI-based approach strongly favors scenarios where invoice layouts vary significantly between vendors and change over time, because AI models (especially deep learning) can generalize and adapt to variations without manual rule updates. Traditional OCR+rule-based pipelines struggle with such variability.

Exam trap

Candidates may think AI is always better, but the question asks for a characteristic that strongly favors AI; limited data or fixed layouts actually favor rule-based, so the trap is selecting those.

How to eliminate wrong answers

Option B is wrong because real-time sub-second latency can be achieved by both approaches; it does not strongly favor AI. Option C is wrong because limited labeled data favors traditional OCR+rule-based or few-shot learning, not standard AI-based models that require large datasets. Option D is wrong because a fixed standardized layout favors rule-based systems, which can be simpler and more accurate.

90
MCQhard

A media company fine-tunes a large language model on Azure Machine Learning to generate sports recaps. After deployment, the model occasionally emits statistics that were never in the source game data. The team wants a systematic way to reduce these unsupported claims without retraining the base model. Which approach BEST addresses this?

A.Increase the fine-tuning dataset size by adding more sports articles and repeat the fine-tuning job.
B.Lower the temperature parameter to 0 and rely on greedy decoding to eliminate fabricated statistics.
C.Apply a post-processing regex filter that removes any numeric token not present in the prompt.
D.Implement a retrieval-augmented generation pipeline that retrieves verified game statistics and constrains the model to cite retrieved passages.
AnswerD

Retrieval-augmented generation grounds generation in an external, verifiable corpus of game statistics. By retrieving relevant passages and instructing the model to base its recap only on those passages, unsupported claims are dramatically reduced because the model has authoritative context at inference time. This addresses the root cause without retraining the base model, matching the team's constraint and providing a systematic, auditable mechanism.

Why this answer

Unsupported claims in generated text are best mitigated by grounding generation in a verifiable source. A retrieval-augmented generation pipeline supplies authoritative game statistics at inference time and instructs the model to rely on them, which reduces fabrication without altering the base model's weights. Deterministic decoding, more fine-tuning data, and regex filtering either miss the root cause or violate the no-retraining constraint.

Exam trap

The trap here is treating hallucination as a sampling-temperature problem, when it is fundamentally a grounding problem that persists even under greedy decoding.

91
MCQmedium

A media company is deploying a generative AI assistant that drafts marketing copy. Legal requires that every generated draft be attributable to source material and that the system must not reproduce copyrighted passages verbatim. The team wants to enforce this at generation time rather than only reviewing outputs afterward. Which implementation approach BEST meets these requirements?

A.Retrieve grounding passages from an approved internal corpus, pass them to the model as context, and attach source citations to the generated draft.
B.Add a post-generation plagiarism check that blocks drafts containing long matching n-grams against a public web index.
C.Fine-tune the base model on the company's existing marketing copy and rely on the fine-tuned weights to avoid verbatim reproduction.
D.Increase the model's temperature setting so the assistant paraphrases source material instead of reproducing it.
AnswerA

Retrieval-augmented generation with an approved corpus constrains the model to cite verifiable sources and lets you attach provenance metadata to each draft. Because grounding passages are supplied at inference time, attribution is enforced at generation time, matching the legal requirement, and the internal corpus reduces the risk of reproducing external copyrighted text.

Why this answer

Grounding generation in an approved internal corpus with retrieval and attaching citations satisfies both legal requirements simultaneously: attributability and reduced verbatim reproduction. The other approaches either act after generation, rely on sampling randomness, or use fine-tuning, none of which guarantees source attribution or constrains output to licensed material at inference time.

Exam trap

The trap here is assuming that raising temperature or fine-tuning automatically prevents verbatim reproduction, when neither attaches source attribution or restricts the model to an approved corpus.

92
MCQmedium

A data science team is preparing a dataset for a binary classification model. The dataset has 95% negative class and 5% positive class. Which technique should they apply to avoid biased model predictions?

A.Apply resampling techniques such as SMOTE or random undersampling
B.Normalise all numerical features to a [0,1] range
C.Shuffle the dataset randomly before splitting into train and test sets
D.Remove all rows with missing values
AnswerA

With only 5% positives, a classifier can achieve 95% accuracy by always predicting the majority class. SMOTE synthesises minority-class examples while random undersampling trims the majority class, rebalancing the training distribution so the model learns the positive class rather than defaulting to the majority.

Why this answer

The dataset is severely imbalanced (95% negative vs. 5% positive), which causes classifiers to favor the majority class and produce biased predictions. Resampling techniques such as SMOTE (Synthetic Minority Over-sampling Technique) generate synthetic minority-class samples, while random undersampling reduces majority-class samples, rebalancing the class distribution so the model learns both classes effectively.

Exam trap

AI0-001 often tests whether candidates confuse data preprocessing steps (normalization, shuffling, imputation) with techniques that specifically address class imbalance, so any option that sounds like 'cleaning data' is a distractor.

How to eliminate wrong answers

Option B is wrong because normalizing numerical features to [0,1] only rescales feature magnitudes and has no effect on class imbalance. Option C is wrong because shuffling the dataset before the train/test split only prevents ordering bias; it does not change the 95/5 class ratio. Option D is wrong because removing rows with missing values addresses data quality, not class imbalance, and could even worsen the imbalance if missingness correlates with the minority class.

93
MCQmedium

A chatbot application uses a system prompt to set the assistant's behavior. The developer wants the LLM to output structured JSON for downstream processing. Which technique BEST ensures the output is valid JSON?

A.Set the temperature to 0 to make output deterministic
B.Include a few-shot example showing a JSON output in the prompt
C.Add a chain-of-thought reasoning step before the output
D.Use the LLM's built-in JSON mode (e.g., response_format='json_object')
AnswerD

JSON mode constrains decoding so the model emits syntactically valid JSON, satisfying the downstream parsing constraint. Unlike prompt-only instructions, which the model may ignore, this enforces the output format at generation time, guaranteeing parseable structured responses.

Why this answer

Many LLMs support a JSON mode that constrains output to valid JSON. System prompts can request JSON but may be ignored. Few-shot examples help but are not foolproof.

Chain-of-thought is for reasoning, not formatting.

94
MCQmedium

A team is building a document intelligence application that extracts key fields from invoices. They have 10,000 labeled invoices. What is the first step in the AI project lifecycle?

A.Model selection – choose a pre-trained vision transformer
B.Data preparation – clean and normalize the invoice images
C.Problem definition – specify which fields to extract and accuracy targets
D.Data acquisition – collect additional invoices from public sources
AnswerC

Before selecting models or labelling schemas, the team must define the business goal: which invoice fields matter and what accuracy threshold counts as success. This scoping drives later data preparation, training and evaluation decisions throughout the lifecycle.

Why this answer

The first step in any AI project lifecycle is problem definition, which involves clearly specifying the business problem, the desired outcomes, and the success criteria. In this case, the team needs to define which fields to extract from invoices and set accuracy targets before proceeding to data preparation, model selection, or acquisition. Without a clear problem definition, subsequent steps lack direction and measurable goals.

Exam trap

AI0-001 often tests the order of the AI project lifecycle; candidates may jump to data preparation or model selection because they seem more technical, but the exam expects recognition that problem definition is always the first step, as it sets the foundation for all subsequent work.

How to eliminate wrong answers

Option A is wrong because model selection occurs later in the lifecycle, after the problem is defined and data is prepared; choosing a model before understanding the problem can lead to mismatched solutions. Option B is wrong because data preparation is important but comes after problem definition; cleaning and normalizing data without knowing what to extract and what accuracy is needed is premature. Option D is wrong because data acquisition is also subsequent to problem definition; the team already has 10,000 labeled invoices, so acquiring more data is not the first step, and doing so without a clear problem definition may result in irrelevant data collection.

95
MCQmedium

An AI application needs to generate structured JSON output from an LLM. The development team wants to ensure the output always conforms to a specific schema. Which prompt engineering technique is MOST suitable?

A.Few-shot examples showing correct JSON
B.System prompt with JSON schema and a 'respond only with valid JSON' instruction
C.Chain-of-thought prompting
D.Fine-tuning the model on JSON datasets
AnswerB

Embedding the schema in the system prompt and instructing the model to respond only with valid JSON constrains generation at the prompt level, satisfying the requirement that output always conforms to a specific schema. The system role carries persistent, high-priority instructions, making schema adherence more reliable than user-turn guidance alone.

Why this answer

Providing the JSON schema directly in the system prompt, combined with an explicit instruction to respond only with valid JSON, is the most direct and reliable way to constrain an LLM's output format. This technique leverages the model's instruction-following capability and schema awareness without requiring examples or retraining, ensuring strict adherence to the desired structure.

Exam trap

The AI0-001 exam often tests the misconception that few-shot examples alone are sufficient for format control, but the trap here is that without an explicit schema and strict instruction, the model may still produce inconsistent or non-compliant output, especially when the schema is complex or the prompt context shifts.

How to eliminate wrong answers

Option A is wrong because few-shot examples can guide the model but do not guarantee strict schema conformance; the model may still deviate from the schema, especially with complex or nested structures. Option C is wrong because chain-of-thought prompting encourages step-by-step reasoning, which often produces intermediate text or explanations, not a clean JSON output, and can actually increase the risk of malformed JSON. Option D is wrong because fine-tuning on JSON datasets is a resource-intensive process that requires significant data, compute, and time, and is overkill for a task that can be solved with a simple prompt-level constraint; it also does not dynamically adapt to schema changes as easily as a system prompt.

96
MCQmedium

A healthcare provider is deploying an AI model to predict patient readmission risk. The model was trained on historical data that includes a feature indicating whether the patient has diabetes. The provider wants to ensure the model does not discriminate based on this feature. Which technique should be used to detect and mitigate bias related to the diabetes feature?

A.Perform a fairness audit by comparing model performance across groups with and without diabetes.
B.Use a different evaluation metric such as AUC-ROC instead of accuracy to assess model performance.
C.Remove the diabetes feature from the dataset and retrain the model.
D.Increase the model's complexity by adding more layers to capture subtle patterns related to diabetes.
AnswerA

A fairness audit evaluates whether the model's predictions are equitable across subgroups defined by the sensitive feature. By comparing metrics like true positive rate or false positive rate between diabetic and non-diabetic patients, the provider can identify disparate impact. This is a standard method to detect bias and informs mitigation strategies such as reweighting or adversarial debiasing.

Why this answer

A fairness audit systematically compares model outcomes across groups defined by the sensitive attribute, revealing disparities. This is the essential first step to detect bias. Once identified, mitigation techniques like reweighting or adversarial debiasing can be applied.

Removing the feature or changing metrics does not directly address the need to detect and mitigate bias.

Exam trap

The trap here is assuming that removing a sensitive feature eliminates bias, when proxy features can still cause discrimination.

97
MCQeasy

A retail company wants its customer support chatbot to answer questions about current promotions that change weekly. The team has an LLM API but does not want to retrain the model each week. Which implementation approach is MOST appropriate?

A.Deploy a smaller open-source model and periodically replace it with a newly trained version each week.
B.Use retrieval-augmented generation to inject current promotion content into the prompt at query time.
C.Increase the model's context window and paste all historical promotions into every system prompt.
D.Fine-tune the LLM on a weekly export of promotion documents.
AnswerB

Retrieval-augmented generation lets the chatbot pull the latest promotion documents and include them as context without changing model weights. Weekly updates become a content-management task rather than a training task, and the model can cite or ground answers in the retrieved material, which directly satisfies the requirement.

Why this answer

Retrieval-augmented generation separates knowledge from model weights, so weekly promotion changes are handled by updating the retrieval corpus rather than retraining. This keeps the LLM stable while ensuring answers reflect current content. Fine-tuning and prompt-stuffing either violate the no-retraining constraint or scale poorly, making retrieval the most appropriate implementation.

Exam trap

The trap here is treating fine-tuning as the default way to add new knowledge, when frequently changing factual content is better handled by retrieval at inference time.

98
MCQmedium

A media company is deploying an AI service that transcribes customer support calls and then summarizes them for agents. The transcription model runs on-premises and produces text, but the summarization LLM is hosted in a public cloud. Compliance requires that no raw call audio or verbatim transcript ever leaves the company network. Which deployment pattern best satisfies this requirement while still using the cloud LLM?

A.Transcribe in the cloud and store the raw audio in a cloud object store, but encrypt the bucket with a customer-managed key.
B.Use a cloud-hosted transcription model and a cloud-hosted summarization model, and configure a private VPC endpoint between them.
C.Run the transcription on-premises, redact or abstract the transcript locally, and send only the redacted, non-verbatim text to the cloud LLM for summarization.
D.Send the raw audio to the cloud LLM and instruct the model via the system prompt to ignore the audio and only produce a summary.
AnswerC

This pattern keeps raw audio and the verbatim transcript inside the company network while still allowing the cloud LLM to perform summarization on de-identified text. Redaction or abstraction removes the compliance-sensitive content before egress. It is a common hybrid deployment pattern for AI solutions where the heavy or sensitive processing stays on-premises and only sanitized text is sent to a hosted model.

Why this answer

The requirement is a data-egress boundary: raw audio and verbatim transcripts must stay on-premises. The only pattern that respects that boundary while still using a cloud LLM is to perform transcription locally, then redact or abstract the text before sending only sanitized content to the hosted model. Encryption, private endpoints, and prompt instructions do not change where the sensitive data is processed, so they cannot satisfy the constraint.

Exam trap

The trap here is assuming that encryption, private networking, or a system prompt can substitute for keeping sensitive data inside the network boundary.

99
MCQeasy

In prompt engineering, which technique involves providing a few correct input-output examples in the prompt to guide the model's response?

A.System prompt engineering
B.Chain-of-thought prompting
C.Few-shot prompting
D.Zero-shot prompting
AnswerC

Few-shot prompting supplies a small number of worked input-output pairs within the prompt itself, letting the model infer the desired pattern and format. This differs from zero-shot, which gives instructions only, and from fine-tuning, which adjusts model weights.

Why this answer

Few-shot prompting provides a small number of input-output examples directly in the prompt so the model can infer the desired task format and pattern without any weight updates. This is distinct from zero-shot (no examples) and chain-of-thought (which elicits step-by-step reasoning rather than demonstrating examples).

Exam trap

AI0-001 often tests the confusion between few-shot (examples in the prompt) and chain-of-thought (step-by-step reasoning), so candidates must read whether the question emphasizes 'examples' or 'reasoning steps'.

How to eliminate wrong answers

Option A is wrong because system prompt engineering refers to setting the model's role, tone, or constraints via a system message, not supplying labeled examples. Option B is wrong because chain-of-thought prompting asks the model to reason step by step (often with 'Let's think step by step'), which is a reasoning technique rather than an example-based one. Option D is wrong because zero-shot prompting provides no examples at all — it relies solely on the instruction.

100
MCQeasy

A team is deploying an anomaly detection system for real-time monitoring of server metrics. The system should alert when metrics deviate significantly from normal patterns. Which type of AI model is MOST suitable?

A.Autoencoder neural network
B.Recommendation system model
C.Linear regression model
D.Image classification model
AnswerA

Autoencoders learn to reconstruct normal input; anomalies yield high reconstruction error, so deviations trigger alerts. This satisfies the real-time, unlabelled server-metric constraint where labelled failure examples are unavailable, unlike supervised classifiers that need pre-tagged anomalies.

Why this answer

An autoencoder neural network is an unsupervised learning model that learns to compress and reconstruct input data. When trained on normal server metrics, it will accurately reconstruct normal patterns but produce high reconstruction error for anomalous patterns, making it ideal for anomaly detection. This allows the system to alert when metrics deviate significantly from normal.

Exam trap

AI0-001 often tests the suitability of different AI models for specific tasks; candidates may choose linear regression because it is simple, but the key is that anomaly detection in complex, high-dimensional data requires models like autoencoders that can capture non-linear patterns and do not require labeled anomalies.

How to eliminate wrong answers

Option B is wrong because recommendation system models are designed to suggest items to users based on preferences, not to detect anomalies in time-series data. Option C is wrong because linear regression models predict a continuous output based on input features and assume a linear relationship; they are not suitable for detecting complex, non-linear anomalies in high-dimensional data. Option D is wrong because image classification models are for categorizing images, not for analyzing numerical server metrics.

101
Multi-Selecthard

A team is developing an AI agent that can answer questions by querying a SQL database and a REST API. The agent should decide which tool to call, parse the response, and reason about the next step. Which THREE concepts should be implemented to build this agent?

Select 3 answers
A.ReAct pattern for iterative reasoning and tool use
B.Content filtering to sanitize database results
C.Function calling to enable the LLM to invoke SQL and API tools
D.Chain-of-thought prompting without tool integration
E.Planning agent that decomposes the question into sub-tasks
AnswersA, C, E

The ReAct pattern interleaves reasoning traces with tool-call actions, so the agent observes each SQL or REST response and reasons toward the next step. This directly satisfies the stem's requirement to decide which tool to call, parse results and reason iteratively.

Why this answer

The ReAct pattern (Reasoning + Acting) enables iterative tool selection. Function calling allows the LLM to output structured tool calls. Planning agents can decompose a question into subtasks.

Chain-of-thought is a reasoning technique but not a full agent framework; content filtering is not needed.

102
MCQmedium

A hospital is implementing an AI triage assistant that suggests urgency levels for emergency department patients. The clinical leadership wants to ensure the system does not systematically undertriage patients from a particular demographic group. Which practice best addresses this requirement during implementation?

A.Report only the overall accuracy of the model across the entire patient population and confirm that it exceeds a pre-agreed target.
B.Remove all demographic features from the training data and assume the resulting model is fair.
C.Evaluate the model's performance separately for each demographic subgroup using metrics such as sensitivity and false negative rate, and remediate disparities before go-live.
D.Require clinicians to override the AI recommendation whenever they disagree, and track the override rate as the fairness measure.
AnswerC

Subgroup evaluation with metrics like sensitivity and false negative rate directly detects systematic undertriage for a demographic group. Undertriage is a false negative in urgency classification, so a higher false negative rate for one group is the signal leadership is worried about. Remediating disparities before deployment, through retraining, reweighting, or threshold adjustment, addresses the requirement at implementation time rather than after harm occurs.

Why this answer

Detecting systematic undertriage requires measuring model performance within each demographic subgroup, especially false negative rates and sensitivity, because aggregate accuracy can conceal group-level harm. Removing demographic features does not remove proxy effects, and override rates reflect clinician behavior rather than model fairness. Subgroup evaluation with remediation before go-live is the practice that directly addresses the clinical leadership's concern.

Exam trap

The trap here is assuming that removing protected attributes from training data makes a model fair, when proxy variables can preserve the disparity.

103
Multi-Selecthard

A company is deploying a large language model (LLM) for internal knowledge management. The model will answer employee questions based on a corpus of confidential documents. The security team requires that the model not leak sensitive information and that responses be accurate. Which TWO techniques should be implemented to meet these requirements? (Choose two.)

Select 2 answers
A.Apply differential privacy during fine-tuning of the LLM on the confidential documents.
B.Implement output filtering to detect and redact sensitive information in the model's responses.
C.Use prompt engineering to instruct the model to refuse answering questions that might reveal sensitive information.
D.Implement retrieval-augmented generation (RAG) with a vector database containing only authorized documents.
E.Deploy the LLM in a sandboxed environment with no internet access and restrict API calls.
AnswersB, D

Output filtering scans generated text for patterns of sensitive data (e.g., PII, confidential terms) and redacts them before delivery. This adds a layer of security by catching leaks that might occur despite other measures. Combined with RAG, it ensures that even if the model inadvertently generates sensitive content, it is not exposed to the user, thus meeting the security requirement.

Why this answer

RAG with an authorized document vector database restricts the model's knowledge to permissible content, enhancing both security and accuracy. Output filtering provides a safety net to redact any sensitive information that might slip through. Together, they address the requirements robustly.

Other techniques like differential privacy or prompt engineering are either insufficient or not directly aimed at preventing leakage in responses.

Exam trap

The trap here is relying on prompt engineering or differential privacy as primary security measures when they do not guarantee prevention of data leakage.

104
Multi-Selectmedium

An AI team is deploying a real-time document intelligence service that extracts key-value pairs from invoices. The pipeline includes an LLM that calls a function to parse structured output. Which TWO testing strategies are essential before production deployment?

Select 2 answers
A.An evaluation framework that compares extracted fields against ground truth for a test set of invoices
B.Load testing to simulate peak invoice volume (e.g., end of month)
C.Regression tests on the model training pipeline to ensure the base LLM hasn't changed
D.Integration tests that call the LLM API with sample invoices and verify the JSON output structure
E.Unit tests for the data pipeline that cleans and normalises invoice images
AnswersA, D

Field-level accuracy is the service's core deliverable, so a ground-truth evaluation framework quantifies extraction precision and recall across a representative invoice test set. It catches systematic errors, such as misread totals or dates, that structural checks alone would pass.

Why this answer

Option A is correct because an evaluation framework that compares extracted key-value pairs against a ground-truth labeled set of invoices is the only way to quantitatively measure field-level accuracy, precision, and recall of the extraction LLM before production, which is essential for a document intelligence service. Option D is correct because integration tests that invoke the actual LLM API with sample invoices and validate the returned JSON structure verify that the function-calling contract, schema conformance, and end-to-end wiring between the pipeline and the model work as expected, catching serialization or tool-call failures that unit tests would miss. Option B is not essential here because load testing addresses throughput and latency at peak volume, which is a performance concern rather than a correctness concern for the extraction quality being validated.

Option C is not essential because regression tests on the base LLM training pipeline are the model provider's responsibility and do not validate this team's deployed inference service. Option E is not essential because unit tests for image cleaning and normalization cover only a preprocessing component, not the LLM extraction behavior or its structured output contract that this deployment hinges on.

Exam trap

The trap is selecting testing strategies that are generally good practices but not essential for this specific use case. Candidates might choose load testing or unit tests for data pipelines because they sound important, but the question asks for essential strategies for the LLM extraction service, which are accuracy evaluation and integration testing.

105
MCQhard

A company is deploying a code generation AI assistant for internal developers. They want to ensure the assistant does not generate code with security vulnerabilities. Which testing approach is MOST critical?

A.Unit tests for the data pipeline that preprocesses prompts
B.Evaluation framework that measures BLEU score on a held-out set of code samples
C.Regression tests that compare outputs of new model versions against a golden dataset
D.Integration tests that send security-focused prompts and validate the generated code against a static analysis tool
AnswerD

Integration tests feed adversarial security prompts to the deployed assistant and pipe generated code through a static analysis tool, catching insecure patterns such as injection or hardcoded secrets. This directly validates the stated requirement that the assistant must not emit vulnerable code, unlike unit or load testing.

Why this answer

The goal is to ensure the assistant does not generate vulnerable code, so the most critical testing is security-focused: send adversarial security prompts and validate the generated code with a static analysis tool (SAST). This directly measures whether the model produces exploitable code and provides actionable feedback, unlike generic quality metrics.

Exam trap

AI0-001 often tests whether candidates confuse general model quality metrics (BLEU, regression tests) with security-specific evaluation, so any option mentioning 'security prompts' plus 'static analysis' is the intended answer.

How to eliminate wrong answers

Option A is wrong because unit tests for the data pipeline verify preprocessing, not the security of generated code. Option B is wrong because BLEU score measures n-gram overlap with reference code and does not detect security vulnerabilities — code can have high BLEU and still be insecure. Option C is wrong because regression tests against a golden dataset check output consistency across model versions, not whether outputs are secure; a golden dataset may not contain security-relevant cases.

106
MCQmedium

A company has an existing AI chatbot that uses a fine-tuned LLM to answer customer queries. They want to add the ability to retrieve real-time order status from their database. Which integration pattern should they use?

A.Implement function calling so the model can trigger a database query and receive the result
B.Prompt the user to check the order status manually
C.Use RAG to retrieve order status from a vector store
D.Embed the database query results directly into the model's training data
AnswerA

Function calling lets the fine-tuned model emit a structured request that the application executes against the database, returning live order status as context for the final answer. This satisfies the real-time retrieval requirement without retraining, unlike embedding-based retrieval, which suits static documents.

Why this answer

Function calling (tool use) lets the LLM emit a structured call to an external function — here, a database query for order status — and then incorporate the returned result into its response. This is the correct pattern for real-time, dynamic data retrieval because the model does not need retraining and the data stays current.

Exam trap

AI0-001 often tests the confusion between RAG (unstructured knowledge retrieval) and function calling (real-time structured data access), so candidates must identify whether the data source is a vector store or a live database.

How to eliminate wrong answers

Option B is wrong because prompting the user to check manually defeats the purpose of an AI assistant and does not integrate the database. Option C is wrong because RAG retrieves from a vector store of embedded documents, which is suited to unstructured knowledge, not real-time transactional queries like order status. Option D is wrong because embedding query results into training data is static, stale immediately, and requires retraining — it cannot provide real-time status.

107
MCQeasy

A marketing team wants to deploy a generative AI assistant that writes product descriptions. Before launch, they must ensure the assistant does not produce copyrighted text or brand-inappropriate claims. Which implementation step best addresses this requirement at generation time?

A.Train the assistant only on the company's own historical product descriptions.
B.Add an output filter that checks generated text for restricted phrases and policy violations before display.
C.Reduce the model's temperature to zero so it always produces the most likely, safest token sequence.
D.Ask users to review each description manually before publishing it.
AnswerB

An output filter inspects the generated text before it reaches users and can block or flag copyrighted snippets, prohibited claims, and brand-inappropriate language. This is a generation-time control that directly enforces the stated requirement and can be updated as policies change. It complements prompt instructions and does not require retraining the model.

Why this answer

The requirement is to prevent non-compliant text from reaching users, which calls for a runtime control on the generated output. An output filter scans for restricted phrases, copyright matches, and policy violations and blocks or flags them before display. Training data, temperature, and manual review do not provide an automated generation-time check that enforces brand and legal policy.

Exam trap

The trap here is assuming that cleaner training data or lower randomness automatically prevents problematic output, when enforcement requires inspecting what the model actually generates.

108
MCQmedium

A financial services company has deployed a credit-risk model that was trained on historical loan data. Regulatory auditors require that the model's decisions be explainable to applicants who are denied credit. The data science team must integrate an explanation capability into the existing production inference pipeline with minimal latency impact. Which approach BEST satisfies the requirement?

A.Log the raw input features and the model's numeric score for each applicant and provide the log file to auditors on request.
B.Publish the model's global feature importance ranking on the company website so applicants can see which factors matter most overall.
C.Apply SHAP (SHapley Additive exPlanations) values to generate per-applicant feature contribution scores at inference time.
D.Replace the production model with a single decision tree so that the full decision path can be shown to each applicant.
AnswerC

SHAP provides locally faithful, additive feature attributions grounded in cooperative game theory, so each denied applicant receives a defensible breakdown of how individual features pushed the decision. It integrates with common model-serving frameworks and can be computed on a per-request basis, which satisfies the audit requirement while keeping explanations tied to the actual deployed model rather than a proxy.

Why this answer

Per-applicant explanations require locally faithful attribution methods that can run inside the inference path. SHAP values attribute the prediction to individual features for that specific applicant, producing the kind of decision rationale regulators expect in adverse action notices. The other approaches either replace the model, provide only aggregate insight, or supply raw data without interpretation, none of which meet the per-decision explainability mandate.

Exam trap

The trap here is confusing global feature importance with local, per-prediction explanations, which are not interchangeable for regulatory adverse action requirements.

109
MCQmedium

A retail bank has deployed a credit-risk scoring model as a REST endpoint behind an API gateway. The model was trained on data from 2019–2023. Compliance now requires the bank to detect when input feature distributions drift away from the training baseline and to trigger retraining before approval rates degrade. Which approach should the bank implement FIRST?

A.Lower the classification threshold from 0.50 to 0.40 so more applicants are approved and the approval rate stays stable.
B.Continuously fine-tune the deployed model on every new loan application that arrives at the endpoint.
C.Enable population stability index (PSI) monitoring on each input feature against the stored training distribution and alert when PSI exceeds a set threshold.
D.Increase the API gateway rate limit and add horizontal replicas so the endpoint can absorb higher application volume.
AnswerC

PSI compares the live feature distribution to the training baseline per feature, which directly surfaces covariate drift in the scoring inputs. Because it is computed on inference payloads, it needs no labels and can trigger retraining before approval-rate degradation becomes visible. This is the standard first-line control for tabular credit models in regulated environments.

Why this answer

Covariate drift is detected by comparing live inference feature distributions with the training baseline, and population stability index is the established metric for tabular models. It operates without ground-truth labels, so it fires early, before approval-rate decay confirms the problem. Fine-tuning on unlabeled traffic, autoscaling, or threshold manipulation all fail to measure distributional change and therefore cannot satisfy the compliance requirement to trigger retraining.

Exam trap

The trap here is assuming that any monitoring on a deployed model, such as latency or throughput dashboards, counts as drift detection when only distribution-comparison metrics actually measure data drift.

110
Multi-Selecthard

A media company is deploying an AI system that generates short news summaries from full articles. Before launch, the responsible AI review board asks the team to define monitoring that will detect harmful or degraded behavior in production. Which TWO monitoring practices should the team implement? (Choose two.)

Select 2 answers
A.Record the GPU utilization of the inference cluster to confirm the model is running within expected compute bounds.
B.Monitor the average token length of generated summaries to ensure outputs stay within the expected range.
C.Track the number of API calls per minute to detect traffic spikes that could indicate abuse or system overload.
D.Track hallucination and factual-consistency rates by automatically comparing generated summaries against source articles using a grounded entailment check.
E.Log and sample generated summaries for human review, with a rubric covering factual accuracy, bias, and tone, and route flagged items to an escalation queue.
AnswersD, E

Automated grounding checks compare each generated claim against the source article and flag unsupported statements, which directly measures hallucination risk in a summarization system. Tracking this rate over time reveals model or data drift, prompt regressions, and changes in source content that degrade factual fidelity. It gives the review board a concrete, ongoing safety signal rather than a one-time evaluation.

Why this answer

Responsible deployment of a generative summarization system requires monitoring both automated factual grounding and structured human review. Grounding checks quantify hallucinations against source articles, while rubric-based human sampling catches bias, tone, and framing issues that metrics miss. Together they provide the safety and quality signals the review board needs.

Infrastructure and length metrics do not measure harmful or degraded content behavior.

Exam trap

The trap here is selecting infrastructure or output-length metrics that feel like monitoring but do not actually detect harmful or factually degraded generated content.

111
MCQmedium

A data scientist is preparing a dataset for a binary classification model to detect fraudulent transactions. The dataset has 1% fraud cases (minority class) and 99% non-fraud cases. Which data preparation technique is MOST appropriate to address the class imbalance before training?

A.Duplicate the minority class samples until the class ratio is 50:50
B.Normalize all features to a range of 0 to 1
C.Random undersample the majority class to match the minority class size
D.Apply SMOTE (Synthetic Minority Over-sampling Technique) to the minority class
AnswerD

SMOTE generates synthetic minority-class samples by interpolating between existing fraud cases, rebalancing the 1:99 ratio without discarding legitimate non-fraud data. This addresses the class imbalance constraint directly, giving the classifier enough minority examples to learn fraud patterns.

Why this answer

SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic samples of the minority class by interpolating between existing minority instances, which addresses class imbalance without simply duplicating data. This reduces overfitting compared to random oversampling and is the most appropriate technique for a 1% fraud detection dataset.

Exam trap

AI0-001 often tests whether candidates confuse data preprocessing steps (normalization) with class imbalance techniques, and whether they know that simple duplication causes overfitting while SMOTE creates synthetic diversity.

How to eliminate wrong answers

Option A is wrong because duplicating minority samples leads to overfitting — the model memorizes exact copies rather than learning generalizable patterns. Option B is wrong because normalization is a feature scaling technique, not a class imbalance remedy; it does not change the class distribution. Option C is wrong because random undersampling discards 99% of majority data, losing valuable information and potentially degrading model performance on legitimate transactions.

112
MCQmedium

A team is implementing a document intelligence solution to extract key-value pairs from invoices. They plan to use a pre-trained vision-language model with a RAG pipeline that indexes invoice images. Which chunking strategy is BEST suited for invoice documents that have a consistent layout but vary in length?

A.Semantic chunking based on sentence boundaries
B.Hierarchical chunking that groups lines into logical sections (header, line items, totals)
C.Fixed-size chunking with 512 tokens per chunk
D.No chunking; pass the entire invoice as one document per query
AnswerB

Hierarchical chunking preserves invoice structure by grouping lines into header, line items, and totals, so retrieval returns coherent key-value pairs rather than fragments. This satisfies the stem's constraint of consistent layout with variable length, where fixed-size splitting would sever field-label relationships.

Why this answer

Invoices typically have sections (header, line items, totals). Hierarchical chunking preserves this structure, enabling retrieval at the section level. Fixed-size may split important fields, semantic chunking is less predictable on structured documents.

113
MCQhard

A logistics company has an AI model that predicts delivery delays. The model performs well in offline evaluation, but after deployment the operations team notices that predictions for a specific region are consistently biased low. The region recently changed its address format in the source system. Which action should the team take to resolve the issue?

A.Remove the affected region from the model's scope and handle its predictions manually.
B.Inspect and update the feature engineering pipeline so the new address format is parsed and encoded consistently with training data.
C.Retrain the model on all historical data, including records with the new address format, without changing the pipeline.
D.Add a post-processing correction that multiplies predictions for the affected region by a fixed factor.
AnswerB

The scenario points to a data pipeline change as the cause of the regional bias. Aligning feature extraction with the training-time format restores the input distribution the model expects, fixing the root cause rather than masking it. This approach also prevents similar issues when other regions change formats.

Why this answer

A change in the source system's address format alters the features the model receives, creating a mismatch with training data and biased predictions. Correcting the feature engineering pipeline restores consistency and addresses the root cause. Post-processing corrections, blind retraining, and excluding the region either mask the issue or reduce coverage without fixing the data defect.

Exam trap

The trap here is treating the regional bias as a model problem and retraining, when the actual defect is an upstream data format change that corrupts features.

114
MCQeasy

A team is building a recommendation system for an e-commerce platform. They need to update recommendations in real-time as users browse. Which integration pattern is MOST suitable?

A.Streaming responses from a single monolithic model
B.Batch processing with nightly updates
C.Async processing queue with delayed responses
D.AI microservices with a REST API
AnswerD

A REST API lets the application call the recommendation model per user interaction, returning fresh ranked results within the request cycle. This request-response pattern satisfies the real-time update constraint, unlike batch scoring, which would serve stale recommendations between scheduled runs.

Why this answer

AI microservices with a REST API allow the recommendation engine to be decomposed into independent, scalable services that can be invoked synchronously on each user interaction, returning fresh recommendations in real time. REST provides a lightweight, stateless request/response contract that fits low-latency, per-request inference, and microservices let the recommendation model scale horizontally and be updated independently of the rest of the e-commerce platform. This combination directly satisfies the 'real-time as users browse' requirement without coupling the model to a monolith or introducing batch/async delays.

Exam trap

AI0-001 often tests the misconception that 'streaming' or 'async' automatically means real-time, when in fact real-time per-request recommendations require a synchronous, low-latency integration pattern such as microservices with a REST API.

How to eliminate wrong answers

Option A is wrong because a single monolithic model creates a tight coupling and a single scaling/update bottleneck, and 'streaming responses' does not by itself provide the per-request, low-latency recommendation updates the scenario requires. Option B is wrong because nightly batch processing updates recommendations only once per day, which is fundamentally incompatible with real-time updates as users browse. Option C is wrong because an async processing queue with delayed responses introduces latency and decouples the response from the user's current session, so recommendations would not reflect the user's live browsing behavior.

115
MCQeasy

Which similarity search metric is BEST for comparing dense vector embeddings when the magnitude of the vectors is not important, only the direction?

A.Euclidean distance
B.Dot product
C.Manhattan distance
D.Cosine similarity
AnswerD

Cosine similarity measures the angle between vectors by computing the dot product of their normalised forms, so it ignores magnitude entirely and reflects only directional alignment. This directly satisfies the stem's constraint that vector length is unimportant, making it the best metric for comparing dense embeddings where orientation, not scale, carries semantic meaning.

Why this answer

Cosine similarity measures the cosine of the angle between two vectors, focusing only on their direction and ignoring magnitude. It is ideal for comparing dense vector embeddings when the magnitude is not important, as it normalizes the vectors. This makes it the best choice for tasks like text similarity where the length of the document should not affect the similarity score.

Exam trap

AI0-001 often tests the confusion between cosine similarity and dot product, especially when vectors are not normalized, so candidates must remember that cosine ignores magnitude.

How to eliminate wrong answers

Option A is wrong because Euclidean distance measures the straight-line distance between vectors and is sensitive to magnitude, so it is not ideal when magnitude is unimportant. Option B is wrong because dot product is influenced by vector magnitude; larger vectors yield higher dot products even if direction is similar. Option C is wrong because Manhattan distance (L1 norm) also depends on magnitude and does not normalize for direction.

116
Multi-Selectmedium

A logistics company is deploying a computer vision model that reads container identification numbers from photos taken at warehouse gates. The model performs well in testing but struggles in production because lighting, camera angles, and container wear vary widely. The team wants to improve robustness before full rollout. Which TWO actions should they take? (Choose two.)

Select 2 answers
A.Collect and label a sample of real gate images, then add them to the training or validation set.
B.Retrain the model using only the highest-resolution images available in the existing dataset.
C.Augment the training set with images that vary brightness, contrast, rotation, and blur to mimic gate conditions.
D.Lower the confidence threshold so the model returns a prediction for every gate image.
E.Increase the model's parameter count by switching to a larger backbone architecture.
AnswersA, C

Real production images capture the true distribution of lighting, angle, and wear that synthetic augmentation only approximates. Adding them to training or validation closes the domain gap and gives the team an honest measure of field performance. This is essential before committing to full rollout.

Why this answer

Robustness gaps between lab and field are closed primarily with representative data. Augmentation simulates the variability of gate conditions, and real labeled gate images supply the actual distribution the model must handle. Together they improve generalization and provide a truthful validation signal.

Larger models, lower thresholds, and resolution filtering do not address the missing data diversity.

Exam trap

The trap here is assuming a bigger model or a looser threshold can compensate for training data that does not represent production conditions.

117
MCQeasy

Which similarity measure is commonly used in vector search to find the angle between vectors, making it well-suited for high-dimensional embeddings?

A.Manhattan distance
B.Euclidean distance
C.Dot product
D.Cosine similarity
AnswerD

Cosine similarity measures the cosine of the angle between two vectors, ignoring magnitude and reflecting only directional alignment. This makes it well-suited to high-dimensional embeddings, where orientation encodes semantic meaning and vector length is largely irrelevant.

Why this answer

Cosine similarity measures the cosine of the angle between two vectors, making it ideal for high-dimensional embeddings because it focuses on orientation rather than magnitude. It is widely used in vector search and NLP tasks where vector direction encodes semantic meaning.

Exam trap

AI0-001 often tests whether candidates confuse dot product with cosine similarity — remember that cosine similarity normalizes by magnitude, while dot product does not, making cosine the angle-based measure.

How to eliminate wrong answers

Option A is wrong because Manhattan distance (L1) measures absolute differences along axes and is sensitive to magnitude, not angle. Option B is wrong because Euclidean distance (L2) measures straight-line distance and is also magnitude-sensitive, which can distort similarity in high-dimensional spaces. Option C is wrong because dot product is related to cosine similarity but is not normalized — it conflates magnitude with angle, so it is not purely a measure of angle.

118
Multi-Selecteasy

A team is building an AI-powered recommendation system for an e-commerce platform. They want to test the system before deployment. Which TWO types of testing are MOST relevant for this AI system? (Select TWO)

Select 2 answers
A.Load testing the web server
B.Integration tests for API calls
C.Evaluation frameworks for model output quality
D.Unit tests for data pipelines
E.Regression testing on the UI
AnswersC, D

Evaluation frameworks assess recommendation output quality, such as relevance, ranking accuracy and coverage, before deployment. This directly satisfies the stem's requirement to test the AI system's behaviour, catching quality issues that conventional functional testing would miss.

Why this answer

Option C (Evaluation frameworks for model output quality) is correct because AI/ML systems produce probabilistic outputs that cannot be validated by traditional assertions; frameworks such as accuracy, precision/recall, F1, BLEU, or ROUGE metrics are needed to assess whether the recommendation model generates relevant, high-quality predictions before deployment. Option D (Unit tests for data pipelines) is correct because the recommendation system depends on feature engineering and ETL/data ingestion pipelines; unit tests verify transformations, schema conformance, null handling, and data integrity so that garbage data does not silently degrade model behavior. Option A (Load testing the web server) is not selected because it measures infrastructure throughput and concurrency rather than the AI system's model correctness or data quality.

Option B (Integration tests for API calls) is not selected because, while useful generally, it tests service-to-service communication rather than the AI-specific concerns of model output quality and data pipeline correctness. Option E (Regression testing on the UI) is not selected because it validates front-end behavior and visual consistency, which is peripheral to testing the AI model and its data pipeline.

Exam trap

AI0-001 often tests whether candidates default to traditional software testing types (load, integration, UI) instead of recognizing AI-specific testing needs like model evaluation and data pipeline validation.

119
MCQhard

A developer is building a RAG system and needs to choose a similarity metric for retrieving document chunks. The embedding model they use produces normalized vectors (unit vectors). Which similarity metric is equivalent to cosine similarity in this case?

A.Jaccard similarity
B.Euclidean distance
C.Manhattan distance
D.Dot product
AnswerD

For unit-length vectors, the dot product equals the cosine of the angle between them, since both magnitudes are 1. Cosine similarity therefore reduces exactly to the dot product, making it the equivalent metric and computationally cheaper because no normalisation division is needed.

Why this answer

For normalized (unit-length) vectors, the dot product equals the cosine of the angle between them, because the magnitudes are both 1. Cosine similarity is defined as the dot product divided by the product of magnitudes; when magnitudes are 1, the denominator is 1, so dot product and cosine similarity are mathematically identical. This makes dot product the correct equivalent metric.

Exam trap

AI0-001 often tests the mathematical equivalence between dot product and cosine similarity for normalized vectors, trapping candidates who assume Euclidean distance is the natural equivalent because both measure 'closeness'.

How to eliminate wrong answers

Option A is wrong because Jaccard similarity measures set overlap (intersection over union) and is not defined for continuous embedding vectors. Option B is wrong because Euclidean distance measures straight-line distance and is not equivalent to cosine similarity even for normalized vectors — it is monotonically related but not equal. Option C is wrong because Manhattan distance (L1) sums absolute coordinate differences and has no equivalence to cosine similarity.

120
MCQeasy

Which component in a RAG system is responsible for converting document chunks into numerical representations that enable similarity search?

A.Vector store index
B.Document chunker
C.Large language model (LLM)
D.Embedding model
AnswerD

Embedding models transform each document chunk into a dense vector, capturing semantic meaning so that similarity search can compare vectors via distance metrics. This satisfies the stem's requirement for numerical representations enabling retrieval, distinct from the LLM that generates answers or the vector store that indexes them.

Why this answer

The embedding model converts text chunks into dense numerical vectors (embeddings) that capture semantic meaning, enabling similarity search in the vector store. It is the component that performs the transformation from raw text to vector representations. Without the embedding model, the vector store would have nothing to index or search.

Exam trap

AI0-001 often tests the distinction between the embedding model (creates vectors) and the vector store (stores/searches vectors), trapping candidates who conflate storage with representation.

How to eliminate wrong answers

Option A is wrong because the vector store index stores and retrieves vectors but does not create them — it relies on embeddings produced elsewhere. Option B is wrong because the document chunker splits documents into smaller pieces but does not convert them to numerical vectors. Option C is wrong because the LLM generates natural language answers from retrieved context; it does not produce the embeddings used for retrieval (unless a separate embedding model is used).

121
MCQmedium

A financial services company is deploying a credit-scoring model built with the AI+ toolkit. The model must produce an explanation for each decision that regulators can review, showing which input features most influenced the score. The data science team has already trained a gradient-boosted tree ensemble. Which approach should the team use to satisfy the regulatory requirement?

A.Publish the model's raw prediction probabilities alongside the input feature values for each applicant.
B.Apply SHAP (SHapley Additive exPlanations) values to the trained ensemble to generate per-instance feature attributions.
C.Replace the ensemble with a single decision tree limited to a depth of three so that the decision path is human-readable.
D.Compute global feature importance by averaging the ensemble's split gains across the entire training set.
AnswerB

SHAP assigns each feature a contribution to the individual prediction based on cooperative game theory, producing locally faithful explanations. For a gradient-boosted ensemble, TreeSHAP computes these values efficiently and exactly, giving regulators a defensible per-decision breakdown. This directly meets the requirement without retraining or replacing the model.

Why this answer

Per-decision explanations require local, instance-level attribution methods. SHAP values quantify how much each feature pushed a single prediction above or below the baseline, and TreeSHAP makes this tractable for tree ensembles. Global importance, model substitution, and raw input-output disclosure all fail to explain why a specific applicant received a specific score.

Exam trap

The trap here is assuming that global feature importance or a simpler surrogate model satisfies a per-decision explainability requirement.

122
MCQhard

An AI team is deploying a fine-tuned LLM for a code generation assistant. They need to ensure the model outputs only syntactically valid JSON for integration with downstream systems. Which prompt engineering technique is MOST effective for enforcing structured output?

A.Enable JSON mode in the API call, specifying the desired JSON schema
B.Provide a few-shot example of a valid JSON response in the prompt
C.Include a system prompt that says 'You are a helpful coding assistant.'
D.Use chain-of-thought prompting to have the model reason step-by-step before answering
AnswerA

JSON mode constrains decoding to emit only tokens forming valid JSON, and the supplied schema further restricts keys, types and nesting. This guarantees syntactic validity at generation time, satisfying the downstream integration constraint that raw prompt instructions alone cannot reliably enforce.

Why this answer

Enabling JSON mode in the API call and specifying the desired JSON schema is the most effective technique because it constrains the model's decoding process at the API level, forcing the output to conform to valid JSON structure. This is a hard constraint enforced by the inference engine, not a soft suggestion in the prompt. It eliminates the risk of malformed output that downstream systems cannot parse.

Exam trap

The trap is assuming that prompt-level techniques (few-shot, system prompts, chain-of-thought) can guarantee structured output; the exam tests whether you know that only API-level constrained decoding (JSON mode/schema) enforces syntax deterministically.

How to eliminate wrong answers

Option B is wrong because few-shot examples only nudge the model toward the desired format; they do not guarantee syntactic validity, and the model can still emit prose or malformed JSON. Option C is wrong because a generic system prompt ('You are a helpful coding assistant') provides no structural constraint on output format. Option D is wrong because chain-of-thought prompting improves reasoning quality but does not enforce JSON syntax; it may even add explanatory text that breaks parsing.

123
MCQeasy

In the AI project lifecycle, after a model is trained and evaluated, it is deployed to a production environment. What is the NEXT critical step to ensure the model continues to perform well over time?

A.Collect more training data
B.Archive the model and start a new project
C.Monitoring the model's performance and data drift
D.Re-train the model from scratch
AnswerC

Monitoring performance and data drift detects degradation once the model is live, satisfying the requirement to sustain accuracy over time. Data drift measures changes in input feature distributions; performance monitoring tracks prediction quality against ground truth. Together they trigger retraining before silent failures erode business outcomes.

Why this answer

After deployment, the next critical step is monitoring the model's performance and detecting data drift, because production data distributions change over time and model accuracy degrades silently. Monitoring closes the MLOps loop by feeding real-world performance signals back to the team, enabling timely retraining or rollback. Without monitoring, degradation goes unnoticed until business impact occurs.

Exam trap

The trap is jumping to remediation actions (retrain, collect data) instead of the diagnostic step (monitoring); the exam tests whether you understand that you must detect degradation before you can fix it.

How to eliminate wrong answers

Option A is wrong because collecting more training data is a response to detected drift or poor performance, not the immediate next step after deployment. Option B is wrong because archiving the model and starting a new project abandons the deployed system and ignores the need to maintain it. Option D is wrong because retraining from scratch is a costly remediation action triggered by monitoring insights, not the first step after deployment.

124
MCQhard

An AI developer is building an agent that can book flights and hotels by calling external APIs. The agent needs to decide which API to call and in what order based on user requests. Which pattern is BEST suited for this multi-step reasoning and tool use?

A.Implement a simple Retrieval-Augmented Generation (RAG) pipeline
B.Fine-tune a model to output API call sequences directly
C.Use function calling with a fixed sequence of API calls
D.Apply the ReAct pattern (Reasoning and Acting)
AnswerD

ReAct interleaves chain-of-thought reasoning with tool calls, letting the agent decide which API to invoke and in what order, then feed results back into reasoning. This satisfies the stem's need for multi-step reasoning plus external API orchestration.

Why this answer

The ReAct (Reasoning and Acting) pattern interleaves natural-language reasoning traces with tool/API calls, allowing the agent to decide which API to invoke, observe the result, and reason about the next step. This iterative loop is ideal for multi-step tasks like booking flights and hotels, where the sequence depends on intermediate results (e.g., flight availability before hotel dates). It is the best-suited pattern for dynamic, multi-step tool use.

Exam trap

The trap is confusing RAG (retrieval for knowledge) or fine-tuning (static behavior) with agentic reasoning; the exam tests whether you recognize that multi-step tool use requires an iterative reason-act loop like ReAct.

How to eliminate wrong answers

Option A is wrong because a simple RAG pipeline retrieves documents to augment generation; it does not orchestrate multi-step API calls or decision-making. Option B is wrong because fine-tuning a model to output API sequences directly is brittle — it cannot adapt to dynamic API responses or errors, and it lacks the reasoning loop needed for conditional sequencing. Option C is wrong because a fixed sequence of API calls cannot handle the conditional logic required (e.g., choosing a hotel based on flight arrival time); it is not adaptive.

125
MCQhard

A financial services firm runs an AI model that scores loan applications. Regulators require the firm to explain any adverse decision to an applicant. The model is a gradient-boosted tree with hundreds of features. Which implementation approach best satisfies the explainability requirement without replacing the model?

A.Replace the gradient-boosted tree with a logistic regression so coefficients are directly interpretable.
B.Publish the model's global feature importance ranking so applicants can see which factors matter most overall.
C.Generate per-applicant explanations using SHAP values that attribute the decision to the most influential features.
D.Log the model's raw input features and output score for each application so auditors can review the data.
AnswerC

SHAP values provide consistent, locally accurate feature attributions for individual predictions, which lets the firm explain why a specific applicant was declined. This satisfies the requirement to explain adverse decisions without discarding the high-performing tree model. The attributions can be presented in applicant-facing language and audited by regulators.

Why this answer

Adverse action explanations must be specific to the individual applicant, and SHAP values provide locally accurate attributions for each prediction. This preserves the gradient-boosted tree's performance while producing reasons that can be communicated and audited. Global importance describes the population, logistic regression replaces the model, and raw logs record inputs without explaining the decision.

Exam trap

The trap here is confusing global feature importance with per-applicant explanations, when adverse action notices require reasons tied to the individual decision.

126
MCQhard

A hospital's AI triage assistant summarizes patient notes for clinicians. During post-deployment monitoring, the team notices the model's outputs drift in tone and length after the vendor silently updated the underlying foundation model. The application code did not change. Which action best restores reproducibility and protects against future silent model changes?

A.Increase the monitoring sample rate so the team detects drift faster on the next update.
B.Lower the max tokens parameter so outputs cannot drift in length after the vendor update.
C.Pin the model to a specific dated version or snapshot and re-run the evaluation suite before promoting any new version.
D.Add a post-processing step that rewrites every summary into a fixed template before display.
AnswerC

Pinning a dated model version or snapshot freezes behavior so results are reproducible, and re-running the evaluation suite before promotion catches regressions caused by vendor updates. This treats the foundation model as a versioned dependency, which is the standard way to control silent changes. It restores reproducibility without freezing the application or abandoning monitoring.

Why this answer

Silent foundation-model updates change behavior even when application code is untouched, so the model must be treated as a pinned, versioned dependency. Freezing a dated snapshot and requiring evaluation before promoting any new version restores reproducibility and blocks unevaluated changes from reaching clinicians. Token limits, output templates, and faster monitoring do not control which model version serves traffic.

Exam trap

The trap here is treating the foundation model as a stable service, when in fact vendor updates can change behavior without any application deployment.

127
Multi-Selectmedium

An AI team is evaluating whether to use AI for a customer segmentation task. They have a dataset of customer demographics and purchase history. Which TWO conditions would make AI a better choice than a traditional rule-based approach? (Select two.)

Select 2 answers
A.The segmentation criteria are well-understood and can be expressed in simple if-then rules
B.The data contains complex, non-linear patterns that are not easily captured by rules
C.The business requires the model to adapt automatically as new customer data arrives
D.The segmentation must be fully explainable to regulators
E.The team has no access to labeled data
AnswersB, C

Why this answer

AI techniques like neural networks or gradient-boosted trees excel at capturing complex, non-linear interactions in high-dimensional data (e.g., purchase sequences combined with demographics) that rule-based systems cannot express without an explosion of brittle, hand-crafted conditions. This makes AI the better choice when the underlying patterns are not linearly separable or easily codified as if-then logic.

Exam trap

The AI0-001 exam often tests the misconception that AI is always superior to rule-based systems, but the trap here is that candidates overlook the specific constraints of explainability (Option D) and data requirements (Option E) that make rule-based approaches more appropriate in those contexts.

128
MCQmedium

A team is implementing a RAG system for legal document retrieval. The documents are long and cover multiple topics. Which chunking strategy is MOST appropriate to ensure each chunk contains coherent information?

A.Hierarchical chunking with overlapping windows
B.Semantic chunking based on topic boundaries
C.Fixed-size chunking with 512 tokens
D.Character-level chunking with no overlap
AnswerB

Semantic chunking splits where embedding similarity between adjacent sentences drops, so boundaries fall at genuine topic shifts rather than arbitrary token counts. Long, multi-topic legal documents therefore yield chunks that each stay within one coherent subject.

Why this answer

Semantic chunking based on topic boundaries is the most appropriate strategy because legal documents are long and cover multiple topics. By splitting at natural topic shifts (e.g., clauses, sections, or argument transitions), each chunk preserves coherent meaning, which is critical for accurate retrieval and generation in a RAG system. This approach avoids mixing unrelated content within a single chunk, which would degrade the quality of retrieved context.

Exam trap

CompTIA often tests the misconception that fixed-size token chunking is always optimal for simplicity, but in domain-specific RAG systems with long, multi-topic documents, semantic boundaries are essential to maintain chunk coherence and retrieval accuracy.

How to eliminate wrong answers

Option A is wrong because hierarchical chunking with overlapping windows adds complexity and redundancy without guaranteeing topic coherence; overlapping windows can introduce duplicate or fragmented information across chunks, which is inefficient for retrieval. Option C is wrong because fixed-size chunking with 512 tokens ignores semantic boundaries, often splitting a single legal argument or clause across two chunks, leading to incomplete or misleading context for the LLM. Option D is wrong because character-level chunking with no overlap destroys all semantic structure, producing arbitrary fragments that are useless for coherent retrieval and generation.

129
MCQeasy

During the data preparation phase of an AI project, a data scientist discovers that the target variable in a binary classification dataset is heavily imbalanced: 95% negative class and 5% positive class. Which technique should be applied to improve model performance on the minority class?

A.Apply oversampling of the minority class using techniques like SMOTE
B.Remove all samples from the majority class to balance the dataset
C.Normalize all numerical features to have zero mean and unit variance
D.Use a train-test split of 80-20 without any modification
AnswerA

SMOTE generates synthetic minority-class samples by interpolating between existing positive instances and their nearest neighbours, rebalancing the 95:5 distribution so the classifier no longer biases towards the majority class. This directly targets the minority-class performance constraint stated in the stem.

Why this answer

The dataset is heavily imbalanced (95% negative, 5% positive), which can cause models to be biased toward the majority class. Oversampling the minority class using techniques like SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic examples of the minority class, balancing the class distribution and improving the model's ability to learn the minority class patterns. This is a standard approach to address class imbalance.

Exam trap

AI0-001 often tests the misconception that accuracy is a good metric for imbalanced data, or that simple removal of majority class is acceptable; candidates might overlook the need for specialized resampling techniques.

How to eliminate wrong answers

Option B is wrong because removing all majority class samples would discard valuable information and likely lead to underfitting and poor generalization. Option C is wrong because normalization is a feature scaling technique that does not address class imbalance. Option D is wrong because a simple train-test split without addressing imbalance will still result in a model that performs poorly on the minority class.

130
MCQhard

A logistics company runs a vision model on edge devices in warehouses to detect damaged packages on conveyor belts. The model must classify each package within 40 milliseconds, and network connectivity to the cloud is unreliable. During a pilot, engineers notice that accuracy on the edge devices is several points lower than the accuracy measured during cloud-based evaluation on the same test images. Which cause is MOST likely?

A.Quantization applied to compress the model for the edge device reduced numerical precision in the weights and activations.
B.The edge devices have less RAM and slower CPUs than the cloud inference servers used during evaluation.
C.The test images were captured with a different camera model than the ones mounted on the conveyor belts.
D.The cloud evaluation used a batch size of 32 while the edge device processes one image at a time.
AnswerA

Converting a full-precision model to a smaller integer format for edge inference changes the numerical values the network computes, which commonly costs a few points of accuracy. Because the cloud evaluation used the original precision, the gap between the two environments points directly at the compression step. This is the classic source of an edge-versus-cloud accuracy delta on identical inputs.

Why this answer

When the same test images yield lower accuracy on the edge than in the cloud, the difference must come from something that changes the computation itself. Quantization to a lower-precision numeric format alters weights and activations, which is the standard cause of a small accuracy regression in compressed edge models. Camera mismatch would affect both environments equally, and hardware speed or batch size do not change the mathematical output for a given image.

Exam trap

The trap here is blaming hardware constraints or input differences for an accuracy gap, when the gap only appears after the model itself was transformed for edge deployment.

131
MCQmedium

A team is building a recommendation system for an e-commerce platform. They want to use collaborative filtering but have a cold-start problem for new users. Which hybrid approach BEST addresses cold start while leveraging collaborative signals?

A.Apply user clustering based on demographic data and then use collaborative filtering within clusters
B.Use only content-based filtering for all users
C.Use matrix factorization with implicit feedback only
D.Implement a hybrid model that combines content-based features with collaborative filtering via a weighted ensemble
AnswerD

A weighted ensemble blends content-based features, which describe new users from available attributes, with collaborative filtering signals from existing users. This supplies recommendations despite missing interaction history, directly resolving the cold-start constraint while preserving collaborative signal use.

Why this answer

A weighted ensemble hybrid combines content-based features (which can profile a brand-new user from their stated preferences, demographics, or first-item interactions) with collaborative filtering signals (which capture behavioral patterns from similar users). This directly mitigates cold start because the content-based component can generate recommendations before enough interaction data exists for collaborative filtering, while the collaborative component takes over as behavioral data accumulates. Pure collaborative filtering fails for new users because there is no interaction history to compute similarities.

Exam trap

AI0-001 often tests the misconception that clustering or demographic segmentation alone solves cold start, when in fact any collaborative method still requires interaction data that new users lack.

How to eliminate wrong answers

Option A is wrong because clustering by demographics and then applying collaborative filtering still requires interaction data within each cluster — a new user with no ratings cannot be matched to neighbors, so cold start persists. Option B is wrong because using only content-based filtering discards collaborative signals entirely, sacrificing the personalization quality that comes from user-item interaction patterns and failing to 'leverage collaborative signals' as the question requires. Option C is wrong because matrix factorization with implicit feedback still needs user-item interaction events to factorize; a new user with zero interactions produces an undefined or random latent vector, so cold start is not solved.

132
MCQmedium

A team is training a image classification model. They split the dataset into training, validation, and test sets. After training, the model achieves 98% accuracy on the training set but only 72% on the test set. Which step in the AI project lifecycle should the team focus on?

A.Data acquisition – collect more data
B.Model selection – use regularization or reduce model complexity
C.Deployment – re-deploy with a different serving framework
D.Data preparation – check for train/test leakage
AnswerB

Regularisation or reduced complexity directly addresses the 26-point train–test gap, which signals overfitting: the model has memorised training noise rather than learning generalisable features. Penalising large weights or shrinking capacity lowers training accuracy while raising test accuracy, satisfying the stem's requirement to close that generalisation gap.

Why this answer

The 98% training accuracy versus 72% test accuracy is a classic signature of overfitting — the model has memorized the training data rather than learning generalizable patterns. The appropriate lifecycle response is to address model complexity through regularization (L1/L2, dropout, early stopping) or by reducing the number of parameters/layers. This directly targets the generalization gap rather than the data pipeline or deployment.

Exam trap

AI0-001 often tests the difference between overfitting (high train, lower test) and underfitting (low train, low test), and candidates frequently jump to 'collect more data' when the real fix is regularization or reduced model complexity.

How to eliminate wrong answers

Option A is wrong because collecting more data would not fix overfitting if the model is already memorizing the training set — the gap is caused by model capacity, not data volume. Option C is wrong because deployment/serving frameworks have no effect on the train-test accuracy gap; the model's learned weights are the problem, not how they are served. Option D is wrong because train/test leakage would typically inflate test accuracy (making it suspiciously high), not produce a large train-test gap where training is much higher than test.

← PreviousPage 2 of 2 · 132 questions total

Ready to test yourself?

Try a timed practice session using only Implementing AI Solutions questions.