Courseiva

CCNA Aio Implementing Ai Questions

75 of 132 questions · Page 1/2 · Aio Implementing Ai topic · Answers revealed

1
MCQeasy

When implementing a vector store for a RAG system, which similarity search metric is MOST commonly used to find the most relevant document chunks for a given query embedding?

A.Manhattan distance
B.Euclidean distance
C.Dot product
D.Cosine similarity
AnswerD

Cosine similarity measures the angle between query and chunk embeddings, ignoring magnitude, which suits text embeddings where direction encodes semantic meaning. It is the default metric in most vector stores such as Azure AI Search.

Why this answer

Cosine similarity is the most common metric for comparing embedding vectors in RAG because it measures the angle between vectors, which works well for high-dimensional semantic embeddings.

2
MCQhard

A team is implementing a RAG system for legal document retrieval. The documents are long (50-100 pages) with clear section headings. They want to ensure that retrieved chunks are semantically coherent and respect document structure. Which chunking strategy is MOST appropriate?

A.Semantic chunking based on sentence embeddings
B.Fixed-size chunking with 256 tokens and no overlap
C.Recursive character text splitting with chunk size 1000 and chunk overlap 200
D.Hierarchical chunking: first split by sections, then further split each section into fixed-size chunks with overlap
AnswerD

Hierarchical chunking splits first on section headings, preserving the document's logical structure, then subdivides oversized sections with overlap. This keeps chunks semantically coherent and respects structure, which fixed-size or naive splitting would fragment across headings.

Why this answer

Hierarchical chunking preserves document structure by first splitting into sections, then further into chunks, maintaining semantic coherence.

3
MCQmedium

A data scientist is preparing a dataset for a binary classification model. The dataset has 95% majority class and 5% minority class. Which data preparation technique is BEST to address the class imbalance?

A.Min-max normalization of all features
B.Random undersampling of the majority class
C.Removing all minority class samples
D.SMOTE oversampling of the minority class
AnswerD

SMOTE generates synthetic minority-class samples by interpolating between existing minority neighbours, rebalancing the 95:5 split without discarding majority data. This gives the classifier more minority examples to learn from, unlike random undersampling which loses information.

Why this answer

SMOTE (Synthetic Minority Over-sampling Technique) is the best choice because it generates new, synthetic minority class samples by interpolating between existing minority instances and their nearest neighbors, rather than simply duplicating them. This directly addresses severe class imbalance (95:5) by enriching the minority class without discarding valuable majority data. Unlike random oversampling, SMOTE reduces the risk of overfitting to exact copies and helps the model learn a more generalizable decision boundary.

Exam trap

AI0-001 often tests the misconception that any resampling technique is equally valid for class imbalance, but the key is recognizing that SMOTE is preferred for severe imbalance because it synthesizes new minority samples without discarding majority data, unlike random undersampling which loses information.

How to eliminate wrong answers

Option A is wrong because min-max normalization only rescales feature values to a [0,1] range and has no effect on class distribution or imbalance. Option B is wrong because random undersampling discards a large portion of the majority class, which with a 95:5 ratio would leave very few majority samples and cause significant information loss, likely degrading model performance. Option C is wrong because removing all minority class samples eliminates the positive class entirely, making binary classification impossible and destroying the dataset's utility.

4
MCQmedium

An AI agent is designed to book flights by calling an external API. The agent must decide which tool to call based on user input, then generate the correct API parameters. Which pattern is MOST appropriate for this workflow?

A.Chain-of-thought prompting only
B.Zero-shot prompting with JSON mode
C.ReAct pattern with tool descriptions and function calling
D.Simple prompt with no tool descriptions
AnswerC

ReAct interleaves reasoning with tool invocation, while function calling supplies the schema that constrains generated API parameters. Together they let the agent select the correct booking tool from its descriptions and emit valid arguments, satisfying the stem's decision and parameter-generation requirements.

Why this answer

The ReAct (Reasoning + Acting) pattern interleaves reasoning traces with tool calls, allowing the agent to decide which tool to invoke based on user input, observe the result, and iterate. Combined with function calling and tool descriptions, the model can select the correct API and generate structured parameters (e.g., JSON schema-conformant arguments) for the flight booking API. This is the canonical pattern for tool-using agents that must choose among multiple tools and produce correct parameters.

Exam trap

AI0-001 often tests the misconception that chain-of-thought or JSON mode alone enables tool use, when in fact tool selection and parameter generation require the ReAct pattern with tool descriptions and function calling.

How to eliminate wrong answers

Option A is wrong because chain-of-thought prompting only elicits reasoning steps; it does not provide a mechanism for the model to invoke external tools or generate structured API parameters, so the agent cannot actually call the flight API. Option B is wrong because zero-shot prompting with JSON mode can produce JSON output but does not give the model tool descriptions or a selection mechanism — it cannot reliably choose among multiple tools or know their parameter schemas. Option D is wrong because a simple prompt with no tool descriptions gives the model no information about available APIs or their parameters, making correct tool selection and parameter generation essentially impossible.

5
Multi-Selectmedium

An organisation is developing a document intelligence system that extracts information from scanned invoices. Which THREE data preparation steps are critical to ensure high extraction accuracy? (Choose THREE.)

Select 3 answers
A.Cleaning and correcting OCR output
B.Removing punctuations and stopwords
C.Normalising all text to lowercase
D.Annotating bounding boxes and field labels
E.Image preprocessing (e.g., deskewing, binarisation)
AnswersA, D, E

Cleaning and correcting OCR output directly addresses the scanned-invoice constraint: OCR introduces character errors on low-quality scans, and those errors propagate into extraction models. Normalising recognised text before training or inference raises accuracy, since the system's inputs are images rather than typed digital text.

Why this answer

Option A (Cleaning and correcting OCR output) is correct because OCR on scanned invoices inevitably introduces character-level errors, and fixing those errors before feeding text to the extraction model directly improves field-level accuracy. Option D (Annotating bounding boxes and field labels) is correct because supervised document intelligence models need labelled ground truth that ties specific fields (e.g., invoice number, total) to their spatial locations to learn accurate extraction. Option E (Image preprocessing such as deskewing and binarisation) is correct because scanned invoices often suffer from rotation, noise, and uneven lighting, and correcting these at the image level raises OCR quality and downstream extraction accuracy.

Option B (Removing punctuations and stopwords) is not appropriate because invoice fields such as dates, currency amounts, and vendor names rely on punctuation and specific tokens, so removing them would destroy critical information. Option C (Normalising all text to lowercase) is not appropriate because case can carry meaning in invoice data (e.g., currency codes, product identifiers, proper names), and lowercasing everything can reduce extraction fidelity.

Exam trap

AI0-001 often tests the distinction between general NLP preprocessing steps (like stopword removal and lowercasing) and document-specific preprocessing (like OCR correction and image enhancement), causing candidates to incorrectly select B or C as critical for extraction accuracy.

6
MCQmedium

A machine learning engineer is building a recommendation system for an e-commerce platform. The system should suggest products based on user purchase history and browsing behavior. Which model selection is BEST suited for this task?

A.Image classification model (e.g., CNN)
B.Linear regression
C.Random forest classifier
D.Collaborative filtering model (e.g., matrix factorization)
AnswerD

Collaborative filtering exploits the interaction matrix between users and items, learning latent factors from purchase and browsing history to predict unseen preferences. This directly matches the scenario's reliance on behavioural signals rather than item content, making matrix factorisation the best-suited approach for personalised product suggestions.

Why this answer

Collaborative filtering models (e.g., matrix factorization) are effective for recommendation tasks using user-item interaction data. Linear regression is for regression, not recommendation. Image classification is unrelated.

Random forests can be used but are less common for collaborative filtering.

7
Multi-Selectmedium

A developer is building an AI agent that needs to call external tools (e.g., weather API, database) and reason about the results to answer user queries. Which THREE components are essential for implementing this agentic workflow?

Select 3 answers
A.Planning capability (e.g., step-by-step decomposition)
B.ReAct (Reasoning + Acting) loop
C.Fine-tuned domain-specific model
D.A vector store for long-term memory
E.Function calling or tool use interface
AnswersA, B, E

Planning capability decomposes a multi-step query into ordered sub-tasks, letting the agent decide which external tool to invoke and in what sequence. Without step-by-step decomposition, the agent cannot coordinate the weather API and database calls needed to reason over combined results before answering.

Why this answer

Option A (Planning capability) is correct because an agentic workflow requires the model to decompose a complex user query into an ordered sequence of steps, deciding which sub-tasks to execute and in what order before invoking tools. Option B (ReAct loop) is correct because the Reasoning + Acting pattern interleaves thought, action, and observation cycles, letting the agent call a tool, reason over the returned result, and decide the next action until the query is resolved. Option E (Function calling or tool use interface) is correct because the agent must have a structured mechanism (e.g., JSON schema-based function/tool definitions) to invoke the weather API or database and receive machine-readable outputs.

Option C (Fine-tuned domain-specific model) is not essential, since a general-purpose LLM with tool-calling and reasoning prompts can drive the workflow without domain fine-tuning. Option D (Vector store for long-term memory) is not essential either, as retrieval-based long-term memory is an optional enhancement rather than a required component for calling tools and reasoning over their results.

Exam trap

The AI0-001 exam often tests the misconception that fine-tuning or vector stores are mandatory for agentic workflows, when in fact the core requirements are planning, a reasoning-acting loop, and a tool-use interface, all achievable with a base model and prompt engineering.

8
MCQmedium

A company is fine-tuning a large language model using PEFT (Parameter-Efficient Fine-Tuning) to reduce GPU memory usage. They have limited hardware and need to fine-tune a 70B parameter model on a single GPU with 24 GB VRAM. Which technique is MOST suitable?

A.Full fine-tuning with gradient checkpointing
B.QLoRA (Quantization-aware LoRA) with 4-bit quantization
C.Instruction tuning with a smaller 7B model
D.LoRA (Low-Rank Adaptation) alone
AnswerB

QLoRA quantises the frozen base weights to 4-bit NF4 and trains only small LoRA adapters, cutting memory enough to fine-tune a 70B model on a single 24 GB GPU. Plain LoRA or full fine-tuning cannot fit within that VRAM budget.

Why this answer

QLoRA combines quantization (4-bit) and LoRA to fine-tune very large models on limited hardware, achieving significant memory reduction while maintaining performance.

9
MCQhard

A team is fine-tuning a large language model using LoRA. They have limited GPU memory. Which technique can further reduce memory consumption while maintaining similar fine-tuning quality?

A.Fine-tune all layers instead of using LoRA
B.Increase the rank of LoRA adapters
C.Use QLoRA with 4-bit quantization of the base model
D.Use a larger batch size
AnswerC

QLoRA quantises the frozen base weights to 4-bit NF4 while training LoRA adapters in higher precision, cutting GPU memory far below standard LoRA. This satisfies the limited-memory constraint, and because only adapters are trained, fine-tuning quality stays comparable.

Why this answer

QLoRA extends LoRA by quantizing the frozen base model to 4-bit precision (using NF4 quantization) while keeping LoRA adapters in higher precision, dramatically reducing GPU memory for the base weights. This allows fine-tuning of large models on a single consumer GPU with quality comparable to 16-bit LoRA. The 4-bit base model is dequantized on the fly during forward and backward passes, so memory savings come primarily from storing the base weights in 4 bits instead of 16.

Exam trap

The trap is confusing LoRA with QLoRA — candidates may think LoRA alone already minimizes memory, but LoRA still stores the base model in 16-bit; only QLoRA's 4-bit quantization of the base model provides the additional memory reduction.

How to eliminate wrong answers

Option A is wrong because fine-tuning all layers requires storing full gradients and optimizer states for every parameter, which increases memory usage by orders of magnitude compared to LoRA. Option B is wrong because increasing the LoRA rank adds more trainable parameters and optimizer state, increasing memory consumption, not reducing it. Option D is wrong because a larger batch size increases activation memory and gradient memory, making GPU memory pressure worse, not better.

10
MCQeasy

A data science team is preparing a dataset for a supervised learning task. They split the data into training and test sets. The team then normalizes the features using the mean and standard deviation calculated from the entire dataset before splitting. What issue does this introduce?

A.It improves model generalization
B.It introduces train/test leakage
C.It causes the model to overfit the training data
D.It reduces the variance of the features
AnswerB

Computing the mean and standard deviation across the whole dataset lets statistics from the test set influence the scaling applied to training features. Test information therefore leaks into training, producing optimistic evaluation results that will not generalise to unseen data.

Why this answer

Computing normalization statistics (mean and standard deviation) on the entire dataset before splitting means the test set's statistics influence the training data transformation. This is a form of train/test leakage: information from the test set leaks into the training pipeline, producing an optimistically biased estimate of model performance. The correct approach is to fit the scaler on the training set only and apply the same transformation to the test set.

Exam trap

AI0-001 often tests whether candidates recognize that preprocessing steps (scaling, imputation, encoding) must be fit only on training data — many candidates focus on model training and forget that leakage can occur in the data preparation pipeline.

How to eliminate wrong answers

Option A is wrong because leakage does not improve true generalization — it only inflates reported metrics, which can mislead model selection and deployment decisions. Option C is wrong because overfitting refers to the model memorizing training data; leakage is a data preparation flaw independent of model complexity, and it can occur even with simple models. Option D is wrong because normalization does reduce variance in a mathematical sense, but that is a side effect, not the issue introduced by computing statistics on the full dataset — the actual problem is leakage.

11
MCQmedium

A data science team is training an image classification model for a medical imaging application. To prevent data leakage, they must partition the dataset correctly. Which approach ensures that no patient images appear in both training and test sets?

A.Split by patient ID so that all images of a patient go to one set only
B.Shuffle the dataset and take the first 80% for training and last 20% for testing
C.Use k-fold cross-validation without grouping
D.Randomly split all images into training and test sets
AnswerA

Grouping all images from one patient into a single partition prevents the same patient's anatomy appearing in both training and test sets. Patient ID is the grouping key that satisfies the stem's no-overlap constraint, since random image-level splitting would leak patient-specific features.

Why this answer

Data leakage occurs when information from the test set leaks into training. Splitting by patient ID ensures that all images from the same patient are kept together in one partition.

12
MCQeasy

During data preparation for a classification model, the data scientist notices that one class has 95% of the samples and the other has only 5%. Which technique is MOST appropriate to address this imbalance?

A.Shuffle the data randomly before each training epoch
B.Remove the minority class samples entirely
C.Use a larger learning rate to force the model to pay attention to the minority class
D.Apply SMOTE (Synthetic Minority Over-sampling Technique) to generate synthetic samples for the minority class
AnswerD

SMOTE interpolates new minority-class points between existing minority neighbours, directly correcting the 95/5 skew. Unlike random oversampling, it does not merely duplicate rows, reducing overfitting risk. This addresses the stated class imbalance before training the classifier.

Why this answer

SMOTE generates synthetic samples of the minority class by interpolating between existing minority instances in feature space, which balances the class distribution and prevents the model from ignoring the minority class. With a 95/5 split, standard classifiers tend to predict the majority class almost exclusively, achieving high accuracy but poor recall on the minority class. SMOTE is the standard, well-established technique for this exact scenario.

Exam trap

AI0-001 often tests whether candidates confuse data-level techniques (SMOTE, resampling) with algorithm-level techniques (class weights, focal loss) — a larger learning rate is a tempting but incorrect 'make the model care' answer.

How to eliminate wrong answers

Option A is wrong because shuffling data before each epoch only changes sample order and does not alter class distribution, so the imbalance persists. Option B is wrong because removing minority samples eliminates the class entirely, making the problem worse and destroying the model's ability to learn it. Option C is wrong because a larger learning rate affects optimization step size, not class weighting, and can cause instability or divergence rather than fixing imbalance.

13
MCQeasy

Which similarity metric is MOST appropriate for comparing dense vector embeddings in a vector store used for document retrieval, when the embeddings are normalized to unit length?

A.Jaccard similarity
B.Manhattan distance
C.Cosine similarity
D.Euclidean distance
AnswerC

With unit-length embeddings, cosine similarity and dot product give identical rankings, but cosine similarity directly measures the angle between vectors, remaining invariant to magnitude. It satisfies the stem's normalisation constraint and is the standard metric for dense retrieval in vector stores.

Why this answer

Cosine similarity measures the angle between two vectors and is the standard metric for comparing dense embeddings in vector stores. When embeddings are normalized to unit length, cosine similarity is mathematically equivalent to the dot product, making it both efficient and semantically meaningful for document retrieval. It focuses on orientation rather than magnitude, which aligns with how embedding models encode semantic similarity.

Exam trap

AI0-001 often tests the relationship between normalization and similarity metrics — candidates may pick Euclidean distance thinking it is equivalent, but cosine similarity is the conventional and mathematically clean choice for unit-length embeddings.

How to eliminate wrong answers

Option A is wrong because Jaccard similarity compares set overlap (intersection over union) and is used for binary or categorical features, not dense continuous vectors. Option B is wrong because Manhattan distance (L1) measures absolute coordinate differences and is sensitive to vector magnitude, making it less appropriate for normalized embeddings where direction matters more than magnitude. Option D is wrong because Euclidean distance (L2) on normalized vectors is monotonically related to cosine similarity but is less commonly used as the primary retrieval metric and can be less numerically stable in high dimensions; cosine similarity is the conventional choice for embedding stores.

14
Multi-Selectmedium

A data scientist is preparing a dataset for training a customer churn prediction model. To prevent train/test leakage, which TWO practices should be followed? (Select TWO)

Select 2 answers
A.Remove duplicate records only from the test set to ensure uniqueness
B.Shuffle the entire dataset randomly before splitting into train and test sets
C.Split the data chronologically (e.g., use data before a certain date for training, after for testing)
D.Normalize numerical features using statistics computed on the entire dataset before splitting
E.Perform feature selection using only the training data, then apply the same features to the test set
AnswersC, E

Chronological splitting trains on earlier records and tests on later ones, mirroring real deployment where future data is unseen. This prevents temporal leakage, satisfying the constraint that test data must not influence or overlap with training information.

Why this answer

Option C is correct because splitting data chronologically (e.g., training on records before a cutoff date and testing on records after that date) respects the temporal order of observations and prevents future information from leaking into the training set, which is essential for time-dependent churn prediction. Option E is correct because feature selection must be performed using only the training data; if the test set influences which features are selected, information from the test set leaks into model development and produces overly optimistic performance estimates. Option A is incorrect because removing duplicates only from the test set does not prevent leakage and can distort the test distribution; duplicate handling should be consistent and decided before splitting.

Option B is incorrect because random shuffling of the entire dataset before splitting can mix past and future observations, which is especially harmful for temporal churn data and does not by itself prevent leakage. Option D is incorrect because computing normalization statistics on the entire dataset before splitting leaks test-set distribution information into training; normalization statistics must be computed only on the training set and then applied to the test set.

Exam trap

The trap is thinking that shuffling or normalizing on the full dataset is harmless — candidates often pick random shuffle or global normalization, not realizing these leak test set information into training.

15
Multi-Selecthard

A company is deploying an LLM-based chatbot that must output responses in a structured JSON format for downstream processing. Which THREE prompt engineering techniques should the team use to ensure the output is valid and correctly structured? (Select three.)

Select 3 answers
A.Include few-shot examples of correct JSON outputs
B.Set temperature to 0 to increase determinism
C.Enable JSON mode or structured output mode in the model API
D.Define the expected JSON schema in the system prompt
E.Use chain-of-thought prompting to reason before output
AnswersA, C, D

Few-shot examples demonstrate the exact JSON structure, key names and nesting the model must reproduce, anchoring its output distribution to valid syntax. This satisfies the downstream parsing constraint by showing rather than merely describing the required format.

Why this answer

Option A is correct because few-shot examples of correct JSON outputs demonstrate the exact structure, key names, and formatting the model should reproduce, which strongly improves adherence to the desired schema. Option C is correct because enabling JSON mode or structured output mode in the model API constrains generation so the response is syntactically valid JSON, directly preventing malformed output. Option D is correct because defining the expected JSON schema in the system prompt gives the model explicit field names, types, and required structure to follow.

Option B is not among the marked correct answers; while low temperature can improve determinism, it does not by itself guarantee valid or correctly structured JSON. Option E is not marked correct because chain-of-thought reasoning improves problem-solving but does not enforce JSON syntax or schema compliance.

16
Multi-Selectmedium

A logistics company is deploying a computer vision model on Azure to detect damaged packages on a conveyor belt. The model runs on Azure IoT Edge devices at each warehouse and must operate during network outages. The team needs to ensure the deployment behaves correctly under intermittent connectivity. (Choose two.)

Select 2 answers
A.Enable Azure IoT Edge offline capabilities by setting the edgeHub module to store and forward telemetry and by using the device's local message queue for inference results.
B.Increase the IoT Hub tier to S3 and enable message routing to a Service Bus queue to guarantee delivery during outages.
C.Deploy the model to an Azure Kubernetes Service cluster in the cloud and expose it through a private endpoint to each warehouse.
D.Configure the model to call the Azure Machine Learning online endpoint for every frame and cache the responses on the device.
E.Package the model as an Azure IoT Edge module and configure the edge device to run inference locally with the module's desired properties set for offline operation.
AnswersA, E

The edgeHub module implements store-and-forward so that telemetry and inference results are buffered locally and delivered when connectivity returns. This preserves detection events generated during outages and prevents data loss, which is essential for a warehouse that must reconcile package damage records after a network gap. Together with local module execution, it satisfies the offline requirement.

Why this answer

Operating during network outages requires local compute and local buffering. Azure IoT Edge modules run inference on the device itself, and the edgeHub module's store-and-forward behavior preserves telemetry and results until connectivity is restored. Cloud-hosted endpoints, higher IoT Hub tiers, and Azure Kubernetes Service all depend on the network being available, so they cannot satisfy the offline requirement for conveyor-belt damage detection.

Exam trap

The trap here is assuming that a higher IoT Hub tier or a private endpoint provides offline resilience, when resilience actually comes from running modules locally and buffering messages on the device.

17
Multi-Selecthard

A logistics company is deploying an AI model that predicts delivery delays. The model is served through an API used by dispatch software. The operations team wants to detect when the model's input data distribution shifts so they can trigger retraining. Which TWO implementation practices best support ongoing detection of data drift in production? (Choose two.)

Select 2 answers
A.Log the model's input feature values and predictions for each request, with timestamps, so the production distribution can be compared against the training baseline.
B.Retrain the model every night on the most recent day of data and automatically promote the new model if its training loss is lower.
C.Ask dispatch operators to report whenever they believe a predicted delay is wrong, and treat a spike in reports as the drift signal.
D.Monitor only the API's average response latency and error rate, and alert when either exceeds a threshold.
E.Compute a statistical drift metric, such as population stability index or KL divergence, between the current input window and the training distribution on a scheduled basis.
AnswersA, E

Logging input feature values and predictions with timestamps creates the raw material for drift detection. Without captured production inputs, there is no way to compare current data against the training distribution. Timestamps allow the team to detect gradual or sudden shifts and to correlate them with external events. This is a foundational practice for any production drift monitoring implementation and is required before statistical drift tests can be applied.

Why this answer

Detecting data drift in production requires capturing the actual inputs the model receives and periodically comparing their distribution to the training baseline. Logging inputs and predictions with timestamps makes the comparison possible, and a scheduled statistical metric such as population stability index or KL divergence turns that data into an alert. Retraining, latency monitoring, and operator reports are either remediation actions or indirect signals that do not directly measure input distribution change.

Exam trap

The trap here is confusing operational monitoring, such as latency and error rate, or remediation such as retraining, with actual detection of input data drift.

18
MCQhard

A media company runs an AI content moderation pipeline that classifies user uploads into allowed, review, and blocked categories. The team notices that the model's blocked decisions have drifted: content that was previously labeled review is now being blocked, and appeals are rising. Which action should the team take FIRST to diagnose the drift?

A.Immediately retrain the model on the most recent two weeks of moderation decisions.
B.Compare the distribution of input features and predicted labels between the current production window and the training baseline.
C.Interview the moderation reviewers to collect qualitative feedback about recent content.
D.Raise the block threshold so fewer items receive the blocked label.
AnswerB

Drift diagnosis begins with measuring how production data and outputs have shifted relative to the reference distribution. Comparing feature and label distributions reveals whether the change is in the inputs, the decision threshold behavior, or both. This evidence directs the next step, such as retraining or threshold recalibration, instead of guessing.

Why this answer

Drift diagnosis is an evidence-gathering step. Comparing production feature and label distributions against the training baseline isolates whether inputs, outputs, or both have moved. That measurement determines whether the fix is retraining, recalibration, or pipeline repair.

Retraining, threshold changes, and interviews all act before the cause is known and can compound the problem.

Exam trap

The trap here is jumping to retraining or threshold adjustment as a reflex instead of first quantifying the drift.

19
MCQmedium

A developer is integrating an AI microservice that accepts image uploads and returns classification labels. The service must handle spikes of up to 1,000 requests per minute but average 100 requests per minute. Which deployment architecture BEST meets these requirements with cost efficiency?

A.Expose the model via a serverless function (e.g., AWS Lambda) with synchronous invocation
B.Use an async processing queue (e.g., RabbitMQ) with a pool of worker instances that auto-scale based on queue depth
C.Deploy the service as a synchronous REST API on a single always-on VM sized for peak load
D.Stream results directly from the model to the client using WebSockets
AnswerB

Queue-depth-based autoscaling lets worker instances expand only during the 1,000-request spikes and contract back to baseline for the 100-request average, so you pay for capacity actually consumed. RabbitMQ decouples ingestion from classification, absorbing bursts without dropping uploads — satisfying both the throughput ceiling and the cost-efficiency constraint.

Why this answer

An async queue with auto-scaling workers decouples ingestion from processing, absorbs bursts by buffering requests, and scales worker count based on queue depth — so you only pay for capacity during actual load. This matches the 10x peak-to-average ratio cost-effectively. Synchronous designs either over-provision for peak or drop requests during spikes.

Exam trap

AI0-001 often tests whether candidates default to 'serverless = always cheapest' — the trap is missing that synchronous serverless hits concurrency limits under bursty load, while async queue + autoscaling workers is the cost-efficient pattern for spiky workloads.

How to eliminate wrong answers

Option A is wrong because synchronous Lambda invocation ties the client to function execution time and concurrency limits; sustained 1,000 rpm bursts can hit account concurrency caps and cause throttling, and synchronous image classification is a poor fit for long-running inference. Option C is wrong because a single always-on VM sized for peak load wastes ~90% of capacity during average periods and provides no horizontal scalability or fault tolerance. Option D is wrong because WebSockets address real-time bidirectional streaming, not burst absorption — the model still needs a scalable backend, and streaming doesn't solve the 10x spike problem.

20
Multi-Selectmedium

A data science team is developing a churn prediction model. Which TWO data preparation best practices are MOST important to prevent overfitting and ensure generalization?

Select 2 answers
A.Split data into training and test sets before any preprocessing
B.Normalize all features using the entire dataset's statistics
C.Use cross-validation to tune hyperparameters
D.Remove outliers based on the full dataset distribution
E.Encode categorical variables with target encoding on the full dataset
AnswersA, C

Holding back a test set before any preprocessing prevents data leakage, since fitting scalers or imputers on the full dataset lets test statistics influence training. This gives an honest estimate of generalisation, directly satisfying the stem's requirement to prevent overfitting on the churn model.

Why this answer

Option A is correct because splitting into training and test sets before any preprocessing prevents data leakage: statistics such as means, standard deviations, or encodings must be learned only from the training set and then applied to the test set, otherwise the model indirectly sees test data and overfitting goes undetected. Option C is correct because cross-validation (e.g., k-fold) tunes hyperparameters on multiple training/validation splits, giving a more reliable estimate of generalization performance and reducing the chance of selecting hyperparameters that overfit a single validation split. Option B is not appropriate because normalizing with statistics computed from the entire dataset leaks test-set information into training; normalization statistics should come from the training fold only.

Option D is not appropriate because removing outliers based on the full dataset distribution also leaks test information and can distort the true data distribution; outlier handling should be based on training data. Option E is not appropriate because target encoding on the full dataset leaks the target variable into the features, a classic cause of overfitting; target encoding must be fit within training folds, ideally with smoothing or out-of-fold encoding.

Exam trap

AI0-001 often tests data leakage: candidates might think normalizing on the full dataset is fine, but it leaks test information and inflates performance.

21
MCQeasy

In the AI project lifecycle, which phase involves splitting the dataset into training, validation, and test sets while ensuring no data leakage?

A.Data preparation
B.Problem definition
C.Data acquisition
D.Model evaluation
AnswerA

Data preparation covers dataset splitting into training, validation and test subsets, and enforces leakage prevention by fitting transformations only on training data before applying them elsewhere. This satisfies the stem's requirement that the split occur without leakage, which later modelling phases cannot retroactively correct.

Why this answer

Splitting the dataset into training, validation, and test sets is a core data preparation step that must be performed before any model training begins. This phase ensures that data leakage is prevented by keeping the test set completely isolated until final evaluation, which is critical for obtaining an unbiased estimate of model performance. In the AI project lifecycle, data preparation encompasses cleaning, transforming, and partitioning the data, making option A the correct phase.

Exam trap

The trap here is that candidates confuse 'data acquisition' (collecting data) with 'data preparation' (cleaning and splitting), leading them to incorrectly choose option C when the question specifically asks about splitting and leakage prevention.

How to eliminate wrong answers

Option B is wrong because problem definition focuses on identifying business objectives and success criteria, not on technical data partitioning or leakage prevention. Option C is wrong because data acquisition involves collecting raw data from sources (e.g., databases, APIs, sensors) and does not include the splitting or leakage-avoidance steps. Option D is wrong because model evaluation occurs after training and uses the already-split test set to assess performance; it does not involve creating the splits or addressing data leakage.

22
Multi-Selectmedium

A team is deploying an AI microservice for real-time object detection in streaming video. Which TWO integration patterns are most appropriate? (Choose two.)

Select 2 answers
A.Streaming responses for real-time inference
B.Batch processing with nightly jobs
C.Synchronous request-response with long timeouts
D.Monolithic application deployment
E.AI microservice architecture
AnswersA, E

Streaming responses push tokens or detection results incrementally as they are produced, rather than buffering a complete payload. This satisfies the real-time constraint: the video pipeline receives bounding-box output with minimal latency, keeping inference aligned with the live stream.

Why this answer

Option A (Streaming responses for real-time inference) is correct because object detection on live video requires continuous, low-latency output as frames arrive, and streaming responses (e.g., gRPC server-streaming or HTTP chunked/SSE) let the service emit detection results incrementally instead of waiting for a full batch to complete. Option E (AI microservice architecture) is correct because packaging the detection model as an independently deployable microservice allows separate scaling, GPU resource allocation, and model versioning without affecting the rest of the streaming pipeline. Option B (Batch processing with nightly jobs) is wrong because nightly jobs introduce hours of latency, which is incompatible with real-time video analytics.

Option C (Synchronous request-response with long timeouts) is wrong because long timeouts block callers and cannot sustain the continuous frame-by-frame throughput that streaming video demands. Option D (Monolithic application deployment) is wrong because a monolith couples the detection workload to unrelated components, preventing independent scaling and rapid model updates needed for real-time inference.

23
Multi-Selecthard

A data scientist is preparing a dataset for a text classification model. To prevent train/test leakage, which THREE practices should they follow?

Select 3 answers
A.Shuffle the entire dataset before splitting to ensure randomness
B.Use time-based splitting for temporal data
C.Perform train/test split before any data cleaning or normalization
D.Apply feature scaling to the entire dataset before splitting
E.Remove duplicate samples and ensure that no text from the same document appears in both sets
AnswersB, C, E

Time-based splitting assigns earlier records to training and later records to testing, respecting chronological order. For temporal data this prevents future information leaking backwards into training, which random splitting would allow, thereby avoiding inflated evaluation results.

Why this answer

Option B is correct because for temporal data, a time-based split (e.g., training on earlier timestamps and testing on later ones) prevents future information from leaking into the training set, which a random split would allow. Option C is correct because performing the train/test split before any cleaning, normalization, or other preprocessing ensures that statistics and transformations are learned only from the training data and not influenced by the test set. Option E is correct because removing duplicate samples and keeping all text from the same document in a single split prevents identical or near-identical content from appearing in both training and test sets, which would inflate performance estimates.

Option A does not belong because shuffling the entire dataset before splitting is not inherently leakage-preventing and can actually cause leakage with temporal or grouped data. Option D does not belong because applying feature scaling to the entire dataset before splitting leaks test-set statistics (mean, variance) into training, which is a classic preprocessing leakage error.

Exam trap

The AI0-001 exam often tests the misconception that shuffling the entire dataset is always safe, but for temporal data or when duplicates exist, shuffling can introduce leakage by mixing future and past samples or spreading identical text across train and test sets.

24
Multi-Selecteasy

A company wants to use AI to automatically detect anomalies in server log data. The data is time-series and labeled with 'normal' and 'anomaly' for the past year. Which TWO techniques are appropriate for this use case?

Select 2 answers
A.Train an image classification model (CNN) on screenshots of log graphs
B.Use a time-series anomaly detection model (e.g., Isolation Forest with sliding windows)
C.Train a supervised classification model (e.g., XGBoost) on extracted features with the labels
D.Use a code generation model to fix the anomalies automatically
E.Build a recommendation system based on user activity logs
AnswersB, C

Isolation Forest works on numerical features; sliding windows capture temporal patterns.

Why this answer

Option B is correct because the data is time-series log data, and techniques like Isolation Forest applied over sliding windows (or similar time-series anomaly detectors) are designed to capture temporal patterns and flag deviations from normal behavior without requiring the labels, which suits anomaly detection on sequential log streams. Option C is correct because the dataset is labeled with 'normal' and 'anomaly' for a full year, so a supervised classifier such as XGBoost can be trained on extracted features (e.g., counts, rates, error codes, latency statistics) to directly learn the mapping from features to the anomaly label. Option A is not appropriate because converting logs to graph screenshots and using a CNN image classifier discards the underlying time-series structure and numeric log semantics, making it an indirect and lossy approach.

Option D is wrong because code generation models fix code rather than detect anomalies in log data, which is the stated goal. Option E is wrong because a recommendation system based on user activity logs addresses personalization, not anomaly detection in server logs.

Exam trap

The AI0-001 exam often tests the distinction between supervised and unsupervised techniques, and candidates mistakenly choose an unsupervised method (like Isolation Forest) when labeled data is available, or they overlook that both supervised and unsupervised approaches can be valid depending on the data and problem framing.

25
MCQmedium

A company is fine-tuning an LLM for a domain-specific task using LoRA. They have limited GPU memory and need to reduce memory footprint without sacrificing fine-tuning quality. Which approach should they consider?

A.Use QLoRA with 4-bit quantized base model
B.Use a larger batch size to speed up training
C.Fine-tune all layers of the base model
D.Increase the rank of LoRA adapters
AnswerA

QLoRA quantises the frozen base model to 4-bit and trains low-rank adapters, cutting GPU memory substantially while preserving fine-tuning quality. This directly satisfies the stem's limited-memory constraint, unlike standard LoRA, which still loads full-precision base weights.

Why this answer

QLoRA combines 4-bit NormalFloat quantization of the base model with LoRA adapters, drastically reducing GPU memory usage while preserving fine-tuning quality through techniques like double quantization and paged optimizers. This directly addresses the constraint of limited GPU memory without sacrificing the model's ability to learn domain-specific tasks effectively.

Exam trap

CompTIA often tests the misconception that increasing model capacity (e.g., higher LoRA rank or full fine-tuning) always improves quality, when in fact memory-constrained environments require efficient techniques like QLoRA that balance resource usage and performance.

How to eliminate wrong answers

Option B is wrong because increasing batch size increases GPU memory consumption, which is counterproductive when memory is limited. Option C is wrong because fine-tuning all layers of the base model requires full gradient storage and optimizer states for every parameter, dramatically increasing memory footprint and defeating the purpose of memory reduction. Option D is wrong because increasing the rank of LoRA adapters increases the number of trainable parameters and their associated optimizer states, raising memory usage without guaranteeing improved fine-tuning quality.

26
MCQmedium

An AI system uses a pre-trained image classification model to detect defects in manufacturing. The team wants to deploy the model in an edge device with limited GPU memory. Which technique should they consider first?

A.Train the model from scratch using a smaller dataset
B.Apply quantization to reduce model size
C.Use a larger model with more parameters for higher accuracy
D.Increase the batch size to improve throughput
AnswerB

Quantization stores weights in lower-precision formats such as INT8 instead of FP32, cutting memory footprint roughly fourfold with minimal accuracy loss. This directly satisfies the stem's constraint of limited GPU memory on the edge device, letting the pre-trained classifier fit and run without retraining or architectural changes.

Why this answer

Quantization reduces the precision of the model's weights and activations (e.g., from 32-bit floating point to 8-bit integer), which significantly shrinks the model size and memory footprint while often maintaining acceptable accuracy. This is the most direct and effective first step for deploying a pre-trained model on an edge device with limited GPU memory, as it requires no retraining and immediately addresses the memory constraint.

Exam trap

The AI0-001 exam often tests the misconception that increasing batch size or model size improves performance in resource-constrained environments, when in fact these actions increase memory demand and are counterproductive for edge deployment.

How to eliminate wrong answers

Option A is wrong because training from scratch on a smaller dataset would likely result in poor accuracy due to insufficient data and would not leverage the benefits of transfer learning, making it an inefficient and risky first step. Option C is wrong because using a larger model with more parameters would increase memory consumption, directly contradicting the goal of deploying on a device with limited GPU memory. Option D is wrong because increasing the batch size increases memory usage per inference step, which would worsen the memory constraint rather than alleviating it.

27
MCQmedium

A data science team is building a binary classifier to detect fraudulent transactions. The dataset has only 2% fraud cases. Which data preparation technique is MOST critical to address this imbalance?

A.Use one-hot encoding on categorical features
B.Remove outliers from the transaction amounts
C.Apply synthetic minority oversampling (SMOTE) to the training set
D.Normalize all numerical features to have zero mean and unit variance
AnswerC

SMOTE generates synthetic fraud examples by interpolating between existing minority-class neighbours, directly addressing the 2% class imbalance that would otherwise bias the classifier toward the majority class. Critically, it must be applied only to the training set, never the test set, to avoid leaking synthetic data into evaluation.

Why this answer

With only 2% fraud cases, the dataset is severely imbalanced, which can cause the classifier to be biased toward the majority class (non-fraud) and achieve high accuracy without learning to detect fraud. SMOTE (Synthetic Minority Oversampling Technique) addresses this by generating synthetic examples of the minority class (fraud) in the training set, balancing the class distribution and improving the model's ability to generalize to fraud cases. This is the most critical technique among the options because it directly tackles the class imbalance problem, which is the primary challenge in this scenario.

Exam trap

CompTIA AI often tests the misconception that data scaling or encoding is the primary fix for imbalance, when in fact techniques like SMOTE that directly modify the class distribution are required.

How to eliminate wrong answers

Option A is wrong because one-hot encoding is a technique for converting categorical variables into a numerical format, but it does not address class imbalance; it is relevant for feature representation, not for balancing the dataset. Option B is wrong because removing outliers from transaction amounts may discard legitimate high-value transactions or even some fraud cases, potentially worsening the imbalance and losing valuable information; outlier removal is for data cleaning, not for handling class imbalance. Option D is wrong because normalizing numerical features to have zero mean and unit variance is a scaling technique that helps gradient descent converge faster and ensures features contribute equally, but it does not alter the class distribution or mitigate imbalance.

28
Multi-Selecteasy

An AI engineer is selecting a PEFT technique to fine-tune a large language model. Which TWO are examples of PEFT (Parameter-Efficient Fine-Tuning)?

Select 2 answers
A.Instruction tuning on a large dataset
B.LoRA
C.QLoRA
D.Gradient checkpointing
E.Full fine-tuning of all parameters
AnswersB, C

LoRA freezes the pretrained weights and injects trainable low-rank decomposition matrices into each transformer layer, updating only a tiny fraction of parameters. This satisfies the PEFT constraint of drastically reducing trainable parameters and memory during fine-tuning, unlike full fine-tuning which updates every weight.

Why this answer

LoRA (Low-Rank Adaptation) is correct because it freezes the pretrained model weights and injects trainable low-rank decomposition matrices into the transformer layers, updating only a tiny fraction of parameters. QLoRA is correct because it extends LoRA by quantizing the base model to 4-bit (NF4) and adding trainable low-rank adapters, further reducing memory while remaining a parameter-efficient method. Instruction tuning (A) is a training objective/paradigm that typically updates all or many parameters, not a PEFT technique itself.

Gradient checkpointing (D) is a memory-saving trick that recomputes activations during backpropagation and does not reduce the number of trainable parameters. Full fine-tuning (E) updates every model parameter, which is the opposite of parameter-efficient fine-tuning.

Exam trap

The trap is confusing memory-saving techniques (gradient checkpointing) or training objectives (instruction tuning) with PEFT — candidates must recognize that PEFT specifically means training a small subset of parameters, not just using less memory.

29
MCQeasy

Which stage of the AI project lifecycle involves splitting data into training, validation, and test sets?

A.Model evaluation
B.Data acquisition
C.Problem definition
D.Data preparation
AnswerD

Data preparation is the lifecycle stage where raw data is cleaned, transformed, and partitioned into training, validation, and test sets, satisfying the requirement to separate data before model training begins. This splitting prevents data leakage and enables unbiased evaluation, making it the stage that directly addresses the question's constraint.

Why this answer

Data preparation is the stage where raw data is cleaned, transformed, and split into training, validation, and test sets. This split is a core part of preparing data for model training — the training set teaches the model, the validation set tunes hyperparameters, and the test set provides an unbiased final evaluation. It occurs before model training and evaluation, making data preparation the correct stage.

Exam trap

The trap is choosing model evaluation because splitting sounds like an evaluation activity — but the split is performed during data preparation, and evaluation merely consumes the already-split test set.

How to eliminate wrong answers

Option A is wrong because model evaluation is the stage where the trained model is assessed on the test set — the split has already happened by then, so evaluation is downstream of the split. Option B is wrong because data acquisition is about collecting and ingesting raw data from sources; splitting happens after cleaning and transformation, not during acquisition. Option C is wrong because problem definition is the very first stage where business goals and success metrics are established — no data is touched yet, so no splitting occurs.

30
MCQmedium

A developer is building an AI agent that needs to call external APIs (e.g., get weather, send email) based on user requests. Which pattern is BEST for enabling the agent to autonomously decide when to call these APIs?

A.Hard-code the API calls in the agent's logic
B.Use a chain-of-thought prompt to reason about the steps
C.Implement function calling in the LLM to generate structured API calls
D.Use a planning agent with a predefined workflow
AnswerC

Function calling lets the LLM emit structured JSON arguments naming which API to invoke and with what parameters, so the agent itself decides when a call is needed rather than following hard-coded logic. This directly satisfies the stem's autonomy requirement, since tool selection happens at inference time based on the user's request.

Why this answer

Function calling in the LLM allows the model to generate structured API calls (e.g., JSON) based on user requests, enabling the agent to autonomously decide when and which API to call. This pattern is best because it leverages the LLM's reasoning to select appropriate functions and parameters, integrating seamlessly with external systems.

Exam trap

AI0-001 often tests the distinction between reasoning techniques (like chain-of-thought) and action mechanisms (like function calling), and candidates may incorrectly choose planning agents or hard-coded logic as the best pattern for autonomous API calls.

How to eliminate wrong answers

Option A is wrong because hard-coding API calls lacks flexibility and does not allow the agent to autonomously decide based on user input. Option B is wrong because chain-of-thought prompting only reasons about steps but does not produce executable API calls; it's a reasoning technique, not an action mechanism. Option D is wrong because a planning agent with a predefined workflow restricts autonomy to a fixed sequence, not dynamic decision-making based on user requests.

31
MCQeasy

A city transit agency wants an AI system to predict bus arrival times. The agency has three years of historical GPS traces, schedule data, and weather records, but no team experienced in building machine learning models. Leadership asks which engagement model will get a working predictor into operations fastest without permanently expanding headcount. Which approach BEST fits?

A.Use a managed AI platform service that ingests the historical data, trains a forecasting model, and exposes a prediction endpoint the agency can call.
B.Publish the raw GPS and weather datasets as open data and rely on external volunteers to build the predictor.
C.Hire a full in-house machine learning team and build a custom training pipeline from scratch on agency servers.
D.Deploy an off-the-shelf spreadsheet forecasting template and have dispatchers manually enter recent arrival times each morning.
AnswerA

A managed service provides the modeling expertise, infrastructure, and deployment path without requiring the agency to hire a data science team, and it can be operational quickly using the data already collected. The agency retains ownership of the operational integration while the vendor handles training and hosting. This matches the speed and headcount constraints precisely.

Why this answer

The agency's real constraints are speed to production and avoiding permanent headcount growth, not building proprietary modeling capability. A managed AI platform absorbs the training, tuning, and hosting work, turning existing historical data into a callable prediction endpoint. Building an internal team is slow and permanent, crowdsourcing has no delivery guarantee, and spreadsheet templates cannot handle the data volume or update frequency the service requires.

Exam trap

The trap here is equating a working AI capability with owning the model, when a managed service can deliver the operational outcome without in-house ML staff.

32
Multi-Selectmedium

A data science team is preparing a dataset for a binary classification model to detect fraudulent transactions. The dataset has 99% legitimate and 1% fraudulent examples. Which TWO techniques should the team apply to improve model performance on the minority class?

Select 2 answers
A.Use class weights in the loss function
B.Oversample the minority class using SMOTE
C.Undersample the majority class randomly
D.Apply data normalisation (z-score) to all features
E.Randomly shuffle the dataset to prevent train/test leakage
AnswersA, B

Class weights scale the loss contribution of each class inversely to its frequency, so the 1% fraudulent examples penalise errors far more heavily. This shifts the decision boundary toward the minority class without altering the underlying 99:1 data distribution.

Why this answer

Option A is correct because assigning class weights in the loss function (e.g., class_weight='balanced' in scikit-learn or a weighted binary cross-entropy) penalizes misclassification of the 1% fraudulent class more heavily, directly countering the 99:1 imbalance during training. Option B is correct because SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic minority-class samples by interpolating between existing fraudulent examples and their k-nearest neighbors, increasing minority representation and helping the model learn the fraudulent decision boundary. Option C is not selected because random undersampling discards potentially useful majority-class data and can increase variance, making it a less preferred technique here.

Option D is not selected because z-score normalization only rescales feature distributions and does nothing to address class imbalance. Option E is not selected because shuffling prevents ordering bias and train/test leakage but has no effect on the skewed class ratio.

Exam trap

The trap is confusing general data preprocessing steps (normalization, shuffling) with techniques specifically designed to handle class imbalance, leading candidates to select options that do not address the core issue.

33
MCQmedium

A media company is deploying a generative AI assistant to summarize customer support calls. The assistant must produce concise summaries in English, but the call transcripts are in Spanish. The team wants to use a single model that can handle both translation and summarization. Which approach is MOST appropriate?

A.Use a multilingual large language model and prompt it to translate and summarize in one step.
B.Use a Spanish-only language model and prompt it to output English summaries.
C.Fine-tune a monolingual English summarization model on translated Spanish transcripts.
D.Deploy a dedicated machine translation model first, then pass the English output to a separate summarization model.
AnswerA

A multilingual LLM can perform cross-lingual tasks like translation and summarization without separate components. It handles the Spanish input and generates an English summary directly, reducing pipeline complexity and latency. This is efficient because the model's training includes multiple languages, enabling it to understand the source and produce the target output in a single inference pass.

Why this answer

A multilingual LLM can directly process Spanish transcripts and generate English summaries in one step, satisfying the need for a single model that handles both translation and summarization. This reduces pipeline complexity and avoids error propagation from separate components, making it the most efficient and effective solution for the described scenario.

Exam trap

The trap here is assuming that a monolingual model can easily be prompted to output in another language without specific training.

34
MCQhard

During testing a chatbot, the QA team observes that the bot sometimes responds with harmful content when given adversarial prompts. Which type of testing should be prioritised to catch these edge cases?

A.Red-teaming and adversarial testing
B.Unit tests for data pipeline functions
C.Regression testing on previously fixed bugs
D.Integration tests for API connectivity
AnswerA

Red-teaming deliberately probes a system with adversarial inputs to expose harmful or unsafe outputs. It targets exactly the edge cases described, where crafted prompts bypass safeguards, making it the testing type that surfaces these vulnerabilities before deployment.

Why this answer

Red-teaming and adversarial testing are specifically designed to probe an AI system for vulnerabilities, including generating harmful or unsafe outputs from adversarial prompts. This approach simulates real-world attacks to uncover edge cases that standard functional tests miss, making it the correct priority for catching harmful content in a chatbot.

Exam trap

The AI0-001 exam often tests the distinction between functional testing (unit, regression, integration) and security-focused testing (red-teaming), trapping candidates who confuse general software testing with AI-specific adversarial evaluation.

How to eliminate wrong answers

Option B is wrong because unit tests for data pipeline functions verify data integrity and transformation logic, not the chatbot's response to malicious inputs. Option C is wrong because regression testing ensures previously fixed bugs remain resolved, but it does not proactively discover new adversarial vulnerabilities. Option D is wrong because integration tests for API connectivity check whether system components communicate correctly, not whether the chatbot produces harmful content under attack.

35
MCQhard

An AI practitioner is fine-tuning a large language model for a domain-specific task using a small labeled dataset (500 examples). They have limited GPU memory. Which technique is MOST suitable?

A.Full fine-tuning of all model parameters
B.QLoRA (Quantized Low-Rank Adaptation)
C.Instruction tuning with the full dataset
D.Retrieval-Augmented Generation (RAG) without fine-tuning
AnswerB

QLoRA quantises the frozen base weights to 4-bit and trains small low-rank adapters, slashing GPU memory far below full fine-tuning. With only 500 labelled examples and constrained VRAM, this parameter-efficient method fits the domain task without exhausting memory.

Why this answer

QLoRA (Quantized Low-Rank Adaptation) is the most suitable technique because it combines 4-bit quantization of the base model with low-rank adapter modules, drastically reducing GPU memory usage while still allowing fine-tuning on a small dataset. This approach preserves the model's pre-trained knowledge and avoids catastrophic forgetting, which is critical when only 500 labeled examples are available.

Exam trap

The AI0-001 exam often tests the misconception that 'fine-tuning always means updating all parameters' or that 'RAG alone can replace fine-tuning for domain adaptation,' leading candidates to overlook memory-efficient adapter methods like QLoRA.

How to eliminate wrong answers

Option A is wrong because full fine-tuning updates all model parameters, requiring substantial GPU memory (often >24GB for a 7B model) and risks overfitting on a tiny dataset of 500 examples. Option C is wrong because instruction tuning typically requires a large, diverse dataset of instruction-response pairs (thousands to millions) and does not inherently reduce memory consumption; it is a data-formatting strategy, not a memory-saving technique. Option D is wrong because RAG without fine-tuning does not adapt the model's internal weights to the domain-specific task, so the model cannot learn the specialized patterns or terminology from the small labeled dataset.

36
MCQeasy

An AI system must extract text from scanned invoices and output structured fields (invoice number, date, total amount). Which type of AI application is this?

A.Chatbot/virtual assistant
B.Code generation
C.Image classification/object detection
D.Document intelligence
AnswerD

Document intelligence combines optical character recognition with field extraction models, converting scanned invoice images into structured key-value output such as invoice number, date and total amount. This matches the stem's requirement to extract text from scans and return named structured fields.

Why this answer

Document intelligence (D) is the correct answer because it specifically refers to AI systems that extract, classify, and structure data from documents like invoices, receipts, and forms. This application uses optical character recognition (OCR) combined with natural language processing (NLP) to identify and output structured fields such as invoice number, date, and total amount, which is exactly what the question describes.

Exam trap

The AI0-001 exam often tests the distinction between general image analysis (object detection) and specialized document processing (document intelligence), so candidates may mistakenly choose image classification because they think scanning an invoice is just 'looking at a picture,' but the key is that the system extracts structured text fields, not just identifies objects.

How to eliminate wrong answers

Option A is wrong because a chatbot/virtual assistant is designed for conversational interactions (e.g., answering questions or performing tasks via dialogue), not for extracting structured data from scanned documents. Option B is wrong because code generation focuses on producing programming code from natural language or other inputs, not on processing scanned invoices. Option C is wrong because image classification/object detection identifies objects or categories within an image (e.g., 'this is a cat' or 'there is a car'), but does not extract specific text fields like invoice numbers or amounts from documents.

37
MCQmedium

An AI system for detecting anomalies in manufacturing sensor data uses a model trained on normal operation data only. During monitoring, the model flags many false positives. Which adjustment is MOST likely to reduce false positives?

A.Switch from an autoencoder to a one-class SVM
B.Add synthetic anomalies to the training set and retrain as a supervised classifier
C.Adjust the anomaly detection threshold to be less sensitive (e.g., require a higher reconstruction error)
D.Increase the size of the training dataset with more normal operation data
AnswerC

Raising the reconstruction-error threshold makes the model flag only stronger deviations, so borderline normal sensor readings no longer trigger alerts. This directly reduces the false positives caused by an overly sensitive threshold, satisfying the stem's requirement to cut false alarms during monitoring.

Why this answer

For an unsupervised anomaly detector trained only on normal data, false positives are typically controlled by tuning the decision threshold — e.g., requiring a higher reconstruction error (for autoencoders) or a lower anomaly score before flagging. Raising the threshold makes the model less sensitive, reducing false positives at the cost of potentially missing some true anomalies. This is the most direct, lowest-effort adjustment.

Exam trap

The trap is over-engineering the fix — candidates pick model swaps or synthetic data generation because they sound sophisticated, but the question asks for the MOST likely adjustment to reduce false positives, which is the simple, direct threshold tuning.

How to eliminate wrong answers

Option A is wrong because switching model families (autoencoder to one-class SVM) does not inherently reduce false positives and may introduce new tuning challenges; it is a larger change without guaranteed benefit. Option B is wrong because adding synthetic anomalies and retraining as a supervised classifier changes the problem framing entirely and requires labeled anomaly data, which may not be available or representative — it is not the most likely quick fix. Option D is wrong because adding more normal data does not address the threshold/sensitivity issue; if the model is already over-flagging normal variation, more normal data alone may not shift the decision boundary enough.

38
MCQmedium

A retail bank is rolling out an AI assistant built on Azure AI Foundry to answer customer questions about account policies. The compliance team requires that any response containing financial advice be routed to a human agent and that all interactions be logged for audit. Which combination of capabilities should the developer implement to meet these requirements?

A.Raise the model's temperature so responses are more varied and enable Azure Policy to restrict which users can query the assistant.
B.Use Azure AI Content Safety to filter harmful content and enable diagnostic logging on the Foundry project.
C.Rely on the model's built-in safety system message to refuse advice questions and enable Application Insights for performance metrics.
D.Implement a custom intent classifier that flags advice-related queries, route flagged sessions to a human agent, and write all turns to Azure Monitor Logs with a retention policy.
AnswerD

A custom intent classifier can be trained to recognize financial-advice requests and trigger a handoff to a human agent, satisfying the routing requirement. Writing every conversation turn to Azure Monitor Logs with a defined retention policy creates the durable, queryable audit trail compliance needs. Together these two mechanisms directly address both stated requirements in a way that is verifiable during an audit.

Why this answer

Meeting both requirements demands a detection mechanism for advice-related content and a durable record of every interaction. A custom intent classifier can flag advice requests and trigger a human handoff, while writing conversation turns to Azure Monitor Logs with a retention policy produces an auditable history. Content filtering, temperature changes, and performance metrics do not provide the routing logic or the compliance-grade logging the bank needs.

Exam trap

The trap here is treating content safety filtering as equivalent to business-specific intent routing, when safety categories and financial-advice detection are entirely different classification problems.

39
MCQhard

A company is building a recommendation system for an e-commerce site. They have historical user-item interaction data. Which approach is most appropriate?

A.Use a large language model to generate random product suggestions
B.Use a pre-trained image classification model to recommend visually similar products
C.Deploy a rule-based system that always recommends best-selling items
D.Train a collaborative filtering model on user-item interactions
AnswerD

Collaborative filtering learns latent user and item factors directly from historical user-item interaction data, such as ratings or purchases, to predict unseen preferences. This matches the stem's available data exactly, unlike content-based approaches that would require item metadata.

Why this answer

Collaborative filtering is the standard approach for recommendation systems when historical user-item interaction data (e.g., ratings, purchases, clicks) is available. It leverages patterns across users and items—such as 'users who liked X also liked Y'—to generate personalized recommendations without requiring explicit item features. Training a collaborative filtering model directly on this interaction matrix captures latent preferences and produces relevant suggestions, making it the most appropriate choice for the described scenario.

Exam trap

AI0-001 often tests the misconception that any advanced model (like LLMs or image classifiers) is suitable for recommendation tasks, when the key is matching the model to the available data type—here, user-item interactions.

How to eliminate wrong answers

Option A is wrong because using an LLM to generate random product suggestions ignores the historical interaction data entirely and produces non-personalized, arbitrary outputs with no grounding in user behavior. Option B is wrong because a pre-trained image classification model recommends based on visual similarity, not on user-item interaction patterns; it would require product images and would not leverage the available interaction data, leading to irrelevant recommendations. Option C is wrong because a rule-based system that always recommends best-selling items is static, non-personalized, and fails to use the rich interaction data to tailor recommendations to individual users.

40
MCQhard

A team is deploying a fine-tuned LLM for code generation. They need to ensure the model output is always valid JSON. Which prompt engineering technique should they use?

A.Chain-of-thought prompting
B.Few-shot examples of valid JSON outputs
C.Temperature setting to 0
D.Using a larger model variant
AnswerB

Few-shot prompting places several complete, valid JSON examples in the prompt, so the model infers the required schema, key names and formatting by pattern matching. This constrains generation toward syntactically valid JSON, satisfying the always-valid-JSON requirement more reliably than zero-shot instructions alone.

Why this answer

Few-shot examples of valid JSON outputs condition the model on the exact schema, key names, and formatting it should produce, dramatically increasing the probability of schema-conformant output. Because LLMs are next-token predictors, showing several input-output pairs where the output is valid JSON teaches the pattern in-context without retraining. This is the most reliable prompt-engineering technique for enforcing structural output constraints.

Exam trap

AI0-001 often tests the confusion between techniques that improve reasoning (chain-of-thought) and techniques that constrain output format (few-shot examples, structured output), tempting candidates to pick chain-of-thought for a formatting problem.

How to eliminate wrong answers

Option A is wrong because chain-of-thought improves reasoning quality but does not constrain output format — a model can reason perfectly and still emit prose or malformed JSON. Option C is wrong because temperature 0 makes output deterministic but not necessarily valid JSON; a deterministic wrong format is still wrong. Option D is wrong because a larger model may be more capable but provides no guarantee of JSON validity and increases cost and latency without addressing the formatting requirement.

41
Multi-Selectmedium

A data scientist is fine-tuning a large language model for a domain-specific task using QLoRA. Which TWO statements correctly describe QLoRA's advantages?

Select 2 answers
A.It enables fine-tuning on consumer-grade GPUs by reducing memory requirements
B.It reduces memory usage by quantizing the base model to 4-bit precision
C.It requires more training data than full fine-tuning to achieve comparable accuracy
D.It trains the full model parameters with low precision
E.It increases inference speed compared to the base model
AnswersA, B

QLoRA quantises the frozen base weights to 4-bit NormalFloat, then backpropagates through low-rank adapters, so gradients and optimiser states stay tiny. This slashes VRAM enough to fine-tune large models on consumer-grade GPUs, directly satisfying the stem's memory-reduction constraint.

Why this answer

Option A is correct because QLoRA (Quantized Low-Rank Adaptation) freezes the base model and trains only small low-rank adapter matrices, drastically cutting the memory needed for gradients and optimizer states, which allows fine-tuning of large models on consumer-grade GPUs. Option B is correct because QLoRA quantizes the frozen base model weights to 4-bit precision (typically using the NF4 data type with double quantization), which is the core mechanism that reduces memory usage while preserving performance. Option C is incorrect because QLoRA does not require more training data than full fine-tuning; it typically achieves comparable accuracy with the same or less data by training only a small number of adapter parameters.

Option D is incorrect because QLoRA does not train the full model parameters; it keeps the base model frozen in 4-bit and trains only the low-rank adapters. Option E is incorrect because QLoRA is a fine-tuning technique and does not inherently increase inference speed over the base model; the 4-bit base model may even require dequantization during inference.

Exam trap

AI0-001 often tests the misconception that quantization improves inference speed or that QLoRA trains all parameters — candidates confuse training-time memory savings with runtime performance gains.

42
MCQmedium

A developer is building an AI agent that needs to call external APIs to complete user requests. The agent must decide which API to call based on the user's natural language input. Which technique should the developer use to enable the agent to invoke APIs?

A.Fine-tuning the LLM on API documentation
B.Chain-of-thought reasoning
C.Few-shot prompting with examples of API calls
D.Function calling
AnswerD

Function calling lets the model emit structured JSON specifying which API to invoke and with which arguments, based on the user's natural-language request. The runtime executes that call and returns results, enabling reliable external API invocation without parsing free-form text.

Why this answer

Function calling (also called tool use) is the native mechanism by which modern LLM APIs let the model emit structured JSON describing which function to invoke and with what arguments. The developer defines a schema for each API, and the model decides at runtime which function to call based on the user's natural language input. This is the standard, purpose-built technique for agentic API invocation.

Exam trap

AI0-001 often tests the confusion between prompting techniques (few-shot, chain-of-thought) and native API capabilities (function calling) — candidates pick prompting because it sounds flexible, missing that function calling is the purpose-built mechanism.

How to eliminate wrong answers

Option A is wrong because fine-tuning on API documentation teaches the model facts about APIs but does not give it a reliable runtime mechanism to emit structured call payloads or handle multi-turn tool orchestration. Option B is wrong because chain-of-thought reasoning improves the model's step-by-step logic but does not by itself produce machine-parseable API invocations or route to the correct endpoint. Option C is wrong because few-shot prompting can nudge the model toward a format but is brittle, consumes context, and lacks the schema enforcement and tool-routing guarantees that native function calling provides.

43
MCQhard

A team fine-tunes a 7B parameter LLM using LoRA on a custom instruction dataset. After training, they observe that the model's outputs are only marginally different from the base model. Which is the MOST likely cause?

A.The dataset contained too many examples, overfitting the adapter
B.The base model was too small to benefit from fine-tuning
C.The LoRA rank was set too low (e.g., r=1), limiting the adapter's capacity to learn the task
D.The learning rate was too high, causing the model to diverge
AnswerC

LoRA rank controls the dimensionality of the low-rank update matrices, so r=1 gives the adapter minimal capacity to capture task-specific patterns. The adapter therefore learns too little, leaving outputs close to the frozen base model.

Why this answer

LoRA (Low-Rank Adaptation) injects trainable low-rank matrices into the model's attention layers. The rank r determines the dimension of these matrices and thus the adapter's capacity to capture task-specific patterns. With r=1, the update matrices are extremely low-rank, severely restricting the number of parameters that can be tuned and limiting the model's ability to learn complex instruction-following behavior.

As a result, the fine-tuned model's outputs remain very close to the base model, as observed.

Exam trap

AI0-001 often tests the misconception that a low LoRA rank is sufficient for any task, confusing parameter efficiency with learning capacity, and candidates may incorrectly attribute marginal output differences to dataset size or learning rate instead of the rank's direct impact on adapter expressiveness.

How to eliminate wrong answers

Option A is wrong because too many examples would typically lead to overfitting, which would cause the model to perform well on training data but poorly on unseen data—not marginal differences from the base model. Option B is wrong because a 7B parameter model is sufficiently large to benefit from fine-tuning; model size is not the primary limiting factor here. Option D is wrong because a high learning rate would cause training instability, loss divergence, or degraded performance, not outputs that are only marginally different from the base model.

44
MCQmedium

Which chunking strategy for RAG is MOST appropriate when documents have a natural hierarchical structure (e.g., sections, subsections)?

A.Hierarchical chunking that preserves document structure
B.Fixed-size chunking with no overlap
C.Semantic chunking based on sentence boundaries
D.Random chunking with varying sizes
AnswerA

Hierarchical chunking preserves the document's section and subsection boundaries, embedding each node with its parent context intact. This directly satisfies the stem's requirement for natural hierarchical structure, unlike fixed-size or semantic splitting, which sever headings from their content and degrade retrieval precision during Microsoft Entra ID-secured RAG queries.

Why this answer

Hierarchical chunking is designed to respect the document's inherent structure—such as sections, subsections, and paragraphs—by creating chunks that align with these boundaries. This preserves the logical flow and context, which is crucial for retrieval-augmented generation (RAG) because the retriever can fetch coherent units that match the query's intent. By maintaining the hierarchy, the system can also leverage parent-child relationships to improve retrieval accuracy and generation quality.

Exam trap

AI0-001 often tests the misconception that any chunking strategy works equally well for all documents, but the key is matching the strategy to the document's inherent structure—hierarchical chunking is specifically designed for documents with natural hierarchies.

How to eliminate wrong answers

Option B is wrong because fixed-size chunking ignores document structure, often splitting mid-sentence or mid-section, which breaks semantic coherence and can lead to fragmented context. Option C is wrong because semantic chunking based on sentence boundaries may still split related ideas across chunks if sentences are part of a larger subsection, and it does not inherently preserve hierarchical relationships. Option D is wrong because random chunking with varying sizes destroys any logical organization, making retrieval unreliable and generation prone to irrelevant or disjointed information.

45
Multi-Selectmedium

A media company is deploying a generative AI assistant that drafts marketing copy for regional campaigns. Legal requires that no customer personal data, unreleased product names, or internal pricing appear in generated output, and that every draft be attributable to a source. The team plans to use retrieval-augmented generation over an approved content repository. Which TWO controls should be implemented to satisfy these requirements? (Choose two.)

Select 2 answers
A.Fine-tune the base model on the entire approved content repository so it memorizes the brand voice and product catalog.
B.Raise the model's temperature setting so the assistant produces more varied and creative marketing phrasing.
C.Apply document-level access control and metadata filtering in the retrieval index so the assistant can only retrieve content the requesting user is authorized to see.
D.Require the assistant to return inline citations that map each generated claim back to the specific retrieved chunk and its source document.
E.Store the full prompt and completion pairs in an unencrypted analytics bucket so the marketing team can review trends.
AnswersC, D

Retrieval is the point where sensitive documents enter the prompt, so enforcing the same permissions as the source repository prevents personal data, unreleased names, and pricing from ever reaching the model. Metadata filtering also restricts retrieval to approved campaign assets. This is the primary control because it blocks leakage at the source rather than trying to scrub output afterward.

Why this answer

Both requirements are met at the retrieval layer. Permission-aware retrieval with metadata filtering stops sensitive documents from entering the prompt at all, which is the strongest form of prevention. Inline citations then make every generated claim traceable to an approved source chunk, enabling review and exposing hallucinations.

Raising temperature, storing raw prompt logs insecurely, or fine-tuning on the whole corpus each either increases leakage risk or removes the ability to control and attribute content.

Exam trap

The trap here is reaching for output-side filters or model tuning when the decisive control is preventing unauthorized documents from being retrieved into the prompt in the first place.

46
Multi-Selecthard

A company is deploying a generative AI application that produces structured JSON output for downstream processing. They want to ensure the output is consistently valid JSON and matches a specific schema. Which THREE techniques should they use? (Select THREE)

Select 3 answers
A.Fine-tune the model on a dataset of JSON outputs
B.Increase the temperature parameter to 1.5
C.Provide few-shot examples of the desired output
D.Include a system prompt specifying the expected JSON schema
E.Use JSON mode (structured output) in the API call
AnswersC, D, E

Few-shot examples demonstrate the exact field names, nesting and formatting expected, steering the model toward the target schema. This pattern-matching improves consistency across calls, though it does not guarantee syntactic validity the way constrained decoding does.

Why this answer

Option C is correct because few-shot examples of the desired JSON output condition the model on the exact structure, key names, and formatting expected, which strongly improves schema adherence. Option D is correct because a system prompt that explicitly specifies the expected JSON schema constrains the model's behavior and instructs it to emit only conforming JSON. Option E is correct because JSON mode (structured output) in the API call enforces syntactically valid JSON at the decoding/API layer and, when combined with a schema, validates the response against that schema.

Option A is not required and is costly: fine-tuning on JSON outputs can bias style but does not guarantee schema-valid JSON at inference time. Option B is wrong because increasing temperature to 1.5 raises randomness and makes malformed or schema-violating output more likely, not less.

Exam trap

AI0-001 often tests the misconception that higher temperature or fine-tuning is needed for structured output — candidates miss that JSON mode plus prompting is the standard, low-cost approach, and that temperature should be lowered, not raised.

47
MCQeasy

A hospital is deploying an AI triage assistant that summarizes patient intake notes for emergency department nurses. Before go-live, the clinical informatics team must define a human oversight process that satisfies both safety and regulatory expectations. Which approach is MOST appropriate?

A.Disable the AI assistant whenever the emergency department is at or above 80 percent capacity to reduce clinician distraction.
B.Allow the AI summary to auto-populate the triage record and rely on nurses to correct errors during their normal chart review.
C.Require a licensed clinician to review and approve the AI-generated summary before it is entered into the triage record.
D.Publish the AI assistant's model card on the hospital intranet and instruct nurses to consult it if they have concerns about a summary.
AnswerC

A human-in-the-loop gate places a licensed professional between AI output and the clinical record, which is the expected oversight pattern for high-stakes healthcare decisions. It preserves clinician accountability, creates a clear audit trail of who approved each summary, and allows the organization to detect systematic model errors during review. This matches regulatory expectations for AI used in patient care.

Why this answer

High-stakes clinical use requires a defined human oversight gate where a qualified professional reviews AI output before it influences care. Requiring clinician approval of each summary establishes accountability, creates an audit trail, and enables detection of model errors. The alternatives either bypass review, introduce arbitrary availability rules, or rely on documentation rather than active oversight.

Exam trap

The trap here is treating documentation such as a model card as equivalent to an active human review step, when oversight requires a person approving output before it affects decisions.

48
MCQmedium

A developer is implementing a RAG system and needs to choose a similarity metric for retrieving document chunks. The embedding model produces normalized vectors. Which metric is computationally efficient and equivalent to cosine similarity for normalized vectors?

A.Euclidean distance
B.Hamming distance
C.Manhattan distance
D.Dot product
AnswerD

For unit-length vectors, the dot product equals cosine similarity because the magnitude denominators are both one, eliminating the normalisation division. It satisfies the efficiency constraint by replacing cosine's two norms and division with a single sum of element-wise products, giving identical rankings at lower computational cost.

Why this answer

For normalized (unit-length) vectors, the dot product is mathematically equivalent to cosine similarity because the cosine formula divides by the product of magnitudes, which are both 1. Dot product is also computationally cheaper since it skips the normalization division, making it the efficient choice for RAG retrieval over normalized embeddings.

Exam trap

AI0-001 often tests the equivalence between dot product and cosine similarity only under the normalized-vector condition — candidates who forget the normalization precondition may incorrectly choose Euclidean distance as 'equivalent.'

How to eliminate wrong answers

Option A is wrong because Euclidean distance measures geometric distance and, while monotonically related to cosine similarity for normalized vectors, requires computing square roots and is not equivalent — it ranks in reverse order relative to similarity. Option B is wrong because Hamming distance counts differing positions in equal-length strings and is used for binary/categorical data, not continuous embedding vectors. Option C is wrong because Manhattan distance (L1) sums absolute coordinate differences and is not equivalent to cosine similarity for normalized vectors.

49
MCQhard

During testing of a customer service chatbot, the team notices that the model sometimes generates plausible-sounding but factually incorrect answers about company policies. Which evaluation approach is BEST to systematically detect and quantify this issue?

A.Regression testing comparing old and new model outputs
B.Unit tests on the data pipeline
C.Integration tests for API calls
D.Evaluation framework with faithfulness and answer relevancy metrics on a held-out test set
AnswerD

Faithfulness metrics measure whether each claim in the generated answer is grounded in the retrieved context, while answer relevancy scores how well it addresses the question. Running both over a held-out test set systematically quantifies hallucinated policy statements.

Why this answer

An evaluation framework with faithfulness and answer relevancy metrics on a held-out test set is the best approach because it directly measures whether the model's output is grounded in the provided source (faithfulness) and whether it actually addresses the user's question (answer relevancy). These are the standard RAG/LLM evaluation metrics designed to detect hallucinations and off-topic answers systematically. A held-out test set ensures the measurement is repeatable and quantifiable across model versions.

Exam trap

AI0-001 often tests the confusion between general software testing types (unit, integration, regression) and AI-specific evaluation metrics, tricking candidates into choosing familiar testing terminology over purpose-built LLM evaluation approaches.

How to eliminate wrong answers

Option A is wrong because regression testing only compares old versus new outputs and cannot detect factual incorrectness if both versions hallucinate the same way. Option B is wrong because unit tests on the data pipeline validate data transformations, not the semantic correctness of LLM-generated text. Option C is wrong because integration tests for API calls only verify connectivity and request/response plumbing, not whether the generated content is factually accurate.

50
MCQhard

A team is deploying a generative AI model for a real-time customer-facing application. They need to balance cost and latency. Which deployment strategy is MOST suitable?

A.Monolithic API with serverless functions
B.Edge deployment on user devices
C.Batch processing with synchronous requests
D.AI microservices with streaming responses and async processing queues
AnswerD

Microservices with streaming responses and async queues decouple request handling from model inference, letting tokens stream to users immediately while queued work absorbs bursts. This satisfies the stem's simultaneous cost and latency constraints for a real-time customer-facing workload.

Why this answer

AI microservices with streaming responses and async processing queues decouple inference from the request lifecycle, allowing the system to handle variable loads efficiently while maintaining low latency for real-time interactions. This architecture balances cost by scaling only the necessary components (e.g., GPU-backed inference services) and uses streaming (e.g., Server-Sent Events or WebSockets) to deliver partial results, reducing perceived latency for the customer.

Exam trap

The AI0-001 exam often tests the misconception that serverless functions (Option A) are always the cheapest and fastest option, but they ignore cold-start latency and the overhead of monolithic orchestration in real-time AI workloads.

How to eliminate wrong answers

Option A is wrong because a monolithic API with serverless functions introduces cold-start latency and tight coupling, which is unsuitable for real-time customer-facing applications where consistent sub-second response times are critical. Option B is wrong because edge deployment on user devices requires significant on-device compute resources, model compression, and frequent updates, which increases deployment complexity and cost, and may not be feasible for large generative models. Option C is wrong because batch processing with synchronous requests is designed for high-throughput, non-real-time workloads (e.g., nightly report generation) and would force users to wait for batch completion, violating the real-time requirement.

51
Multi-Selecthard

An organization wants to fine-tune a 7B parameter LLM for a specialized legal document summarization task. They have a small labeled dataset (500 examples) and limited GPU budget. Which THREE techniques should they consider? (Choose three.)

Select 3 answers
A.Use LoRA (Low-Rank Adaptation)
B.Create an instruction-tuning dataset with input-summary pairs
C.Train a new model from scratch on legal text
D.Full fine-tuning of all model parameters
E.Use QLoRA with 4-bit quantization
AnswersA, B, E

LoRA freezes the 7B base weights and trains only small low-rank adapter matrices, cutting trainable parameters and optimiser memory dramatically. This directly satisfies the limited GPU budget while remaining effective with only 500 labelled legal examples.

Why this answer

Option A (LoRA) is correct because Low-Rank Adaptation freezes the pretrained 7B weights and trains small low-rank adapter matrices, drastically reducing trainable parameters and GPU memory, which fits the limited GPU budget. Option B (instruction-tuning dataset with input-summary pairs) is correct because the 500 labeled examples must be formatted as supervised instruction/response pairs so the model learns the specific legal summarization mapping during fine-tuning. Option E (QLoRA with 4-bit quantization) is correct because QLoRA quantizes the frozen base model to 4-bit (e.g., NF4) and trains LoRA adapters, further cutting memory so a 7B model can be fine-tuned on modest GPUs.

Option C is not appropriate because training a new model from scratch on legal text requires massive compute and data, far beyond 500 examples and a limited GPU budget. Option D is not appropriate because full fine-tuning updates all 7B parameters, demanding very high GPU memory and compute that the organization's budget cannot support.

52
MCQmedium

A machine learning engineer is deploying a real-time anomaly detection system for manufacturing sensor data. The system must process thousands of readings per second with minimal latency. Which deployment architecture is BEST suited?

A.Batch processing using Apache Spark jobs triggered hourly
B.Serverless functions deployed on a CDN
C.A monolithic web application with a relational database
D.AI microservices with an async processing queue and streaming responses
AnswerD

Microservices decouple ingestion from inference via an async queue, absorbing thousands of readings per second without blocking, while streaming responses return anomaly results with minimal latency. This satisfies the throughput and latency constraints that synchronous request-response architectures cannot.

Why this answer

Real-time anomaly detection on high-throughput sensor streams requires low-latency, scalable, event-driven processing. An AI microservices architecture with an async processing queue and streaming responses decouples ingestion from inference, allowing horizontal scaling of model-serving instances and back-pressure handling. This design keeps latency low while processing thousands of readings per second, unlike batch or monolithic approaches.

Exam trap

AI0-001 often tests real-time vs. batch trade-offs — candidates pick batch or serverless because they sound scalable, but only an async streaming microservices design meets the low-latency, high-throughput requirement.

How to eliminate wrong answers

Option A is wrong because hourly Spark batch jobs introduce massive latency and cannot support real-time detection. Option B is wrong because serverless functions on a CDN are designed for edge HTTP request handling, not sustained high-throughput stream processing with stateful ML inference. Option C is wrong because a monolithic web app with a relational database creates a bottleneck and cannot elastically scale to thousands of readings per second with minimal latency.

53
Multi-Selectmedium

A company is choosing between fine-tuning and RAG for a legal document assistant. Which TWO factors would MOST strongly favor RAG over fine-tuning?

Select 2 answers
A.The legal documents are updated frequently (weekly)
B.The model needs to understand complex legal terminology
C.The queries require deep reasoning across multiple documents
D.The assistant must cite specific sources for its answers
E.The company has limited compute budget for training
AnswersA, D

Frequent weekly updates suit RAG because retrieval pulls current documents at query time, so the assistant reflects new content without retraining. Fine-tuning bakes knowledge into model weights, requiring repeated, costly retraining to stay current — directly satisfying the stem's freshness constraint.

Why this answer

Option A is correct because RAG retrieves from an external knowledge base at query time, so weekly-updated legal documents can be re-indexed without retraining the model, whereas fine-tuning would require repeated, costly retraining to keep pace. Option D is correct because RAG naturally returns the retrieved passages that grounded the answer, enabling precise source citations, while a fine-tuned model's parametric knowledge offers no traceable provenance. Option B does not favor RAG, since understanding complex legal terminology is largely a function of the base model's pretraining and can be addressed by either approach.

Option C does not favor RAG either, as deep multi-document reasoning is a model capability rather than a retrieval benefit, and RAG alone does not guarantee it. Option E is not a strong RAG advantage, because building and operating a vector store and retriever also incurs infrastructure and compute costs, so a limited training budget does not decisively favor RAG.

Exam trap

AI0-001 often tests the trade-offs between RAG and fine-tuning, where candidates may incorrectly assume fine-tuning is always better for domain-specific terminology or reasoning, overlooking RAG's advantages in dynamism and citation.

54
Multi-Selectmedium

A data scientist is preparing a dataset for a regression model. The dataset contains 100 features, some of which are highly correlated. To improve model performance and reduce overfitting, which TWO techniques should the data scientist apply? (Select TWO)

Select 2 answers
A.Feature selection
B.Dimensionality reduction (e.g., PCA)
C.Data augmentation
D.Adding more hidden layers to the neural network
E.Increasing the learning rate
AnswersA, B

Feature selection removes redundant, highly correlated predictors, directly cutting the 100-feature dimensionality that drives overfitting. By retaining only informative variables, the regression model generalises better, satisfying the stem's requirement to improve performance while reducing overfitting from multicollinearity.

Why this answer

Feature selection (A) is correct because it removes irrelevant or redundant features from the 100-feature set, directly reducing the dimensionality and mitigating overfitting caused by highly correlated predictors. Dimensionality reduction such as PCA (B) is also correct because it transforms the correlated features into a smaller set of uncorrelated principal components, preserving most variance while reducing overfitting and improving model performance. Data augmentation (C) is not appropriate here since it expands training data (typically for images/text) and does not address feature correlation or dimensionality.

Adding more hidden layers (D) increases model complexity and would likely worsen overfitting rather than reduce it. Increasing the learning rate (E) is a training hyperparameter change that does not address correlated features and can even destabilize convergence.

Exam trap

AI0-001 often tests whether candidates confuse techniques that reduce model complexity (feature selection, PCA) with techniques that increase capacity or change training dynamics (more layers, higher learning rate), so read the goal — reducing overfitting — carefully.

55
MCQhard

A team is implementing a RAG system for a large legal document repository. They need to chunk the documents for efficient retrieval. The documents contain long sections with subsections, and the team wants to preserve the hierarchical structure. Which chunking strategy is MOST appropriate?

A.Hierarchical chunking that preserves section and subsection boundaries
B.Overlapping chunking with a 10% token overlap
C.Semantic chunking based on topic segmentation
D.Fixed-size chunking with 512 tokens per chunk
AnswerA

Hierarchical chunking splits on section and subsection boundaries, preserving parent-child relationships so retrieved chunks retain context. Fixed-size or semantic chunking would fragment subsections and lose the legal document's structure, which the stem explicitly requires be preserved.

Why this answer

Hierarchical chunking splits documents along their natural section and subsection boundaries, so each chunk retains its parent heading context and the retrieval system can return the most relevant subsection without losing the document's structural meaning. This is essential for legal documents where a clause's meaning depends on which article and sub-clause it belongs to. Preserving hierarchy also enables parent-child retrieval, where a small child chunk is matched but the larger parent section is fed to the LLM.

Exam trap

The trap here is assuming any chunking strategy that mentions overlap or semantic similarity is automatically superior, when the question's keyword 'hierarchical structure' points specifically to structure-preserving chunking.

How to eliminate wrong answers

Option B is wrong because overlapping chunking with 10% token overlap is a generic technique to avoid cutting sentences mid-thought; it ignores document structure and can split a subsection across chunks, breaking legal context. Option C is wrong because semantic chunking groups text by topic similarity using embeddings, which is useful for unstructured prose but does not guarantee preservation of explicit section/subsection boundaries. Option D is wrong because fixed-size 512-token chunking is the crudest approach, arbitrarily cutting text and frequently severing headings from their content, which is especially damaging for hierarchical legal documents.

56
MCQmedium

A financial services firm is deploying a credit-scoring model built with Amazon SageMaker. Compliance requires that every prediction be explainable to a loan officer and to regulators. The model uses gradient boosting on 120 features. Which approach BEST satisfies the explainability requirement while keeping the production model unchanged?

A.Log the raw feature vector for every request and have analysts manually reconstruct the decision using spreadsheet formulas.
B.Enable SageMaker Model Monitor with a data quality baseline and publish the monitoring reports to the loan officers.
C.Replace the gradient boosting model with a single decision tree so that every prediction follows a human-readable path.
D.Use SageMaker Clarify to generate SHAP-based feature attributions for each inference request through the Clarify explainer endpoint.
AnswerD

SageMaker Clarify supports SHAP-based explanations and can be configured as an online explainer endpoint that returns per-request feature attributions alongside the prediction. This gives loan officers and regulators a ranked contribution of each feature for the individual decision without retraining or replacing the gradient boosting model, directly meeting the explainability requirement while leaving the production model untouched.

Why this answer

Explainability for regulated decisions requires per-prediction attributions tied to the actual model. SageMaker Clarify's SHAP explainer endpoint returns feature-level contributions for each request while the gradient boosting model remains in production, satisfying both accuracy and compliance. Monitoring tools track aggregate drift, simpler models change behavior, and manual reconstruction is not auditable, so only the Clarify-based approach meets the constraint of leaving the production model unchanged.

Exam trap

The trap here is confusing model monitoring for drift with per-prediction explainability, since both are described as making a model 'transparent'.

57
MCQmedium

A retail bank deploys a machine learning model that scores loan applications. Compliance requires that the bank be able to explain to regulators why any individual applicant was denied, in terms of the applicant's own feature values. The model is a gradient-boosted tree ensemble trained on 200 features. Which approach BEST satisfies this requirement?

A.Apply SHAP (SHapley Additive exPlanations) values to produce per-applicant feature attributions for each decision.
B.Log the model's predicted probability alongside the applicant's raw input record for each decision.
C.Report the global feature importance ranking produced by the tree ensemble's gain-based split scores.
D.Retrain the model as a logistic regression and present the learned coefficients as the explanation.
AnswerA

SHAP values come from cooperative game theory and assign each feature a signed contribution to the model's output for that specific instance, so a denial can be explained as a sum of feature-level reasons drawn from the applicant's own data. This is model-agnostic, works for tree ensembles, and directly produces the individualized, feature-value-based justification regulators are requesting.

Why this answer

The requirement is individualized, feature-value-based justification for each denial on a complex ensemble. SHAP attributions decompose a single prediction into additive feature contributions, which is exactly what an adverse-action explanation needs. Global importance, coefficient tables, and raw input logging all describe population behavior or stored data rather than the reason for one applicant's specific outcome, so they cannot satisfy the regulator's question.

Exam trap

The trap here is assuming that any interpretability output, such as a global feature importance chart, counts as an explanation for an individual decision.

58
MCQhard

A recommendation system for an e-commerce platform is experiencing a high false positive rate in its anomaly detection module, causing legitimate transactions to be flagged as fraudulent. The team wants to reduce false positives without significantly increasing false negatives. Which action is MOST effective?

A.Decrease the anomaly detection threshold
B.Increase the anomaly detection threshold
C.Use a different anomaly detection algorithm
D.Increase the size of the training dataset
AnswerB

Raising the anomaly threshold makes the detector flag only higher-scoring cases, so fewer legitimate transactions cross into the positive class. This directly lowers false positives, though some true anomalies may be missed. The stem prioritises reducing false positives without a large false-negative rise.

Why this answer

Increasing the anomaly detection threshold makes the model more conservative about flagging transactions as anomalous, which directly reduces false positives (legitimate transactions incorrectly flagged). Because the threshold is raised, only more extreme deviations trigger an alert, so some true anomalies may be missed, but the question explicitly accepts a small increase in false negatives as a tradeoff.

Exam trap

AI0-001 often tests the direction of the threshold tradeoff, tricking candidates who confuse 'reducing false positives' with 'increasing sensitivity' and therefore choose to decrease the threshold.

How to eliminate wrong answers

Option A is wrong because decreasing the threshold makes the model more sensitive, flagging even mild deviations as anomalies, which would increase false positives rather than reduce them. Option C is wrong because switching algorithms is a larger, less predictable change that does not guarantee a reduction in false positives and may introduce new failure modes. Option D is wrong because increasing training data volume can improve overall model quality but does not directly control the precision/recall tradeoff the way threshold adjustment does.

59
MCQeasy

A data scientist is preparing a dataset for a classification model. The dataset has missing values in several features and features with very different scales. Which two data preparation steps should be applied?

A.Cleaning and normalization
B.Outlier removal and binning
C.Feature selection and dimensionality reduction
D.Data augmentation and one-hot encoding
AnswerA

Cleaning imputes or removes missing values so the classifier receives complete records, while normalisation rescales features onto a comparable range, preventing large-magnitude features from dominating distance or gradient calculations. Both address the dataset's stated defects.

Why this answer

Cleaning handles missing values (e.g., imputation), and normalization scales features to a similar range, which is important for many ML algorithms.

60
MCQeasy

A small marketing team has built an internal AI assistant that answers questions about their product catalog. They want to deploy it quickly with minimal infrastructure management and pay only for what they use. The team has no Kubernetes expertise and wants the provider to handle scaling, patching, and availability. Which deployment option best matches these constraints?

A.Use a serverless function that loads the model on each invocation and returns a response.
B.Use a managed AI platform endpoint that hosts the model, handles scaling and patching, and bills per request or per token.
C.Provision a Kubernetes cluster, deploy the model behind an inference server, and configure a horizontal pod autoscaler.
D.Run the model on a dedicated virtual machine that the team patches and scales manually.
AnswerB

A managed AI platform endpoint matches every constraint: the provider handles scaling, patching, and availability, the team avoids Kubernetes operations, and billing is consumption-based. This is the fastest path to production for a small team without infrastructure expertise. It also provides built-in monitoring and integration with the provider's SDKs, reducing the amount of custom operational code the team must write and maintain.

Why this answer

The team's constraints are minimal infrastructure management, no Kubernetes expertise, and consumption-based billing. A managed AI platform endpoint is designed for exactly this situation: the provider owns scaling, patching, and availability, and the team pays per request or token. Kubernetes, serverless functions with per-invocation model loading, and self-managed virtual machines each impose operational or performance costs that conflict with the stated requirements.

Exam trap

The trap here is treating serverless functions as a universal low-ops answer, when cold starts and memory limits make them a poor fit for interactive model serving.

61
MCQmedium

A company wants to build a conversational agent that can handle complex multi-step tasks such as booking a flight, reserving a hotel, and scheduling a car rental in a single session. The agent must be able to break down the user's request into sub-tasks, call external APIs, and reason about the results. Which design pattern is BEST suited for this requirement?

A.Retrieval-Augmented Generation (RAG) with a vector store
B.An agentic workflow implementing the ReAct pattern with tool use
C.A single large language model prompt with all instructions
D.Fine-tuning a model on a dataset of flight, hotel, and rental conversations
AnswerB

ReAct interleaves reasoning traces with tool calls, letting the agent decompose the request, invoke flight, hotel and car APIs, then reason over each result before the next step. This satisfies the multi-step, external-API constraint that a single prompt or plain chain cannot handle.

Why this answer

The ReAct (Reasoning + Acting) pattern interleaves chain-of-thought reasoning with tool/API calls, allowing the agent to decompose a multi-step request, invoke external services (flight, hotel, car APIs), observe results, and iterate. This is precisely what's needed for orchestrating dependent sub-tasks across multiple systems. Agentic workflows with tool use are the industry-standard design for multi-step task automation with LLMs.

Exam trap

AI0-001 often tests the misconception that RAG or fine-tuning alone can handle multi-step agentic tasks — candidates must distinguish retrieval (knowledge augmentation) and fine-tuning (behavior shaping) from true agentic orchestration with tool use and reasoning loops.

How to eliminate wrong answers

Option A is wrong because RAG with a vector store only augments generation with retrieved context — it does not provide the reasoning loop, tool invocation, or state management needed to execute multi-step bookings. Option C is wrong because a single monolithic prompt cannot dynamically call external APIs, handle intermediate results, or recover from failures across a multi-turn workflow. Option D is wrong because fine-tuning teaches style/domain patterns but does not grant the model the ability to call APIs, plan, or reason iteratively at runtime — it's a static capability, not an orchestration mechanism.

62
MCQmedium

A retail company wants to forecast weekly demand for thousands of SKUs across stores. The data includes strong seasonal patterns, promotional calendars, and intermittent demand for slow-moving items. The team has limited time and wants a baseline before investing in custom deep learning. Which approach is BEST as the initial production model?

A.Apply a clustering algorithm to group SKUs by sales similarity and forecast only the cluster centroids.
B.Train a single global deep neural network on all SKUs and stores with embeddings for product and location.
C.Deploy a large language model to generate demand forecasts by prompting it with recent sales figures and promotion descriptions.
D.Use a classical time-series method such as exponential smoothing or ARIMA applied per SKU-store series with seasonal and promotional regressors.
AnswerD

Classical methods handle seasonality and trend well, can incorporate promotional regressors, and produce per-series forecasts that are easy to explain and monitor. They are fast to implement, require modest compute, and are robust for intermittent demand when configured appropriately. This makes them a strong initial production baseline before considering more complex models.

Why this answer

A fast, reliable baseline for seasonal and intermittent demand is a classical time-series method applied per SKU-store series with promotional regressors. These methods are well understood, quick to deploy, and produce explainable per-series forecasts that support monitoring. Deep learning, LLM prompting, and clustering either require disproportionate investment or do not yield actionable SKU-level forecasts.

Exam trap

The trap here is reaching for deep learning or LLMs by default when classical forecasting methods already handle seasonality and promotions with far less effort.

63
MCQeasy

A data scientist is building a binary classification model to predict customer churn. The dataset has 90% non-churn and 10% churn. After training, the model achieves 90% accuracy, but the recall for the churn class is only 20%. Which metric should the team primarily focus on to evaluate the model's effectiveness?

A.Recall for the churn class
B.Accuracy
C.Area Under the ROC Curve (AUC-ROC)
D.Precision for the non-churn class
AnswerA

Recall for the churn class directly exposes the model's failure to identify actual churners, which accuracy conceals under the 90/10 imbalance. Optimising recall ensures genuine churn cases are captured, satisfying the stem's requirement to evaluate effectiveness on the minority class rather than overall correctness.

Why this answer

With 90% non-churn and 10% churn, a model that predicts 'no churn' for everyone achieves 90% accuracy but 0% churn recall. The current model's 20% churn recall means it misses 80% of actual churners, which is the business-critical failure. Recall for the churn class directly measures how many true churners are caught, making it the primary metric.

Exam trap

AI0-001 often tests the misconception that high accuracy means a good model — candidates must recognize that on imbalanced data, accuracy is misleading and minority-class recall is the meaningful metric.

How to eliminate wrong answers

Option B is wrong because accuracy is misleading on imbalanced data — the majority-class baseline already achieves 90%, so accuracy hides the model's failure on the minority class. Option C is wrong because AUC-ROC can look acceptable even when minority-class recall is poor, since it aggregates performance across all thresholds and is dominated by the majority class. Option D is wrong because precision for the non-churn class is trivially high when the model over-predicts non-churn, and it does not measure the ability to identify churners.

64
MCQmedium

A team is developing an AI agent to assist users with multi-step tasks such as booking a flight, reserving a hotel, and scheduling a car rental. The agent needs to reason about the order of steps and handle dependencies. Which pattern is BEST suited?

A.Simple tool use without reasoning
B.Using a single prompt with all instructions
C.Fine-tuning a model to output all steps at once
D.ReAct pattern (Reasoning and Acting)
AnswerD

ReAct interleaves reasoning traces with tool-calling actions, letting the agent decide each next step based on observed results. This supports sequencing dependent tasks such as flight, hotel, and car rental bookings, where later steps rely on earlier outcomes.

Why this answer

The ReAct pattern (Reasoning and Acting) is best suited because it interleaves reasoning traces with tool calls, allowing the agent to dynamically plan and adjust steps based on intermediate results. For multi-step tasks with dependencies (e.g., booking a flight before a hotel), ReAct enables the agent to reason about order, handle failures, and call external APIs step-by-step, which is essential for robust task completion.

Exam trap

The key trap is that candidates overlook the need for dynamic reasoning and tool interaction, assuming that a single large prompt or fine-tuned output can handle all multi-step tasks. However, only the ReAct pattern provides the necessary step-by-step planning and tool use.

How to eliminate wrong answers

Option A is wrong because simple tool use without reasoning lacks the ability to plan or handle dependencies; it can only execute isolated function calls without context. Option B is wrong because using a single prompt with all instructions cannot adapt to dynamic changes or intermediate results; it assumes a static plan that fails if any step requires conditional logic or error recovery. Option C is wrong because fine-tuning a model to output all steps at once (single-shot generation) cannot handle real-time feedback from external systems or adapt to variable execution order, making it brittle for interactive multi-step workflows.

65
MCQhard

A hospital's radiology AI triage model was validated at 94% sensitivity on a curated research dataset. After six months in production, clinicians report that it misses many positive cases on images from a newly installed scanner. The data science team confirms the model has not been retrained. Which action should the team take FIRST to diagnose and correct the problem?

A.Increase the model's classification threshold so that more cases are flagged as positive.
B.Compare the statistical distribution of production images from the new scanner against the training data to detect covariate shift.
C.Fine-tune the model on the most recent production images that clinicians have labeled as missed positives.
D.Replace the triage model with a larger architecture trained from scratch on a combined dataset of all available scanners.
AnswerB

The model was validated on one data distribution and now receives images from different hardware, which strongly suggests covariate shift: the input features moved while the learned relationship stayed fixed. Measuring distribution differences such as intensity histograms, resolution, and noise characteristics identifies whether inputs are out of the training envelope, which is the root cause to establish before choosing any remediation.

Why this answer

Performance collapse on images from new hardware, with no retraining, points to a change in the input distribution rather than a change in the label relationship. Detecting and quantifying that shift through distribution comparison establishes the root cause before any intervention. Threshold changes, targeted fine-tuning, and full retraining are all remedies that presuppose a cause and can waste effort or introduce new bias if applied blindly.

Exam trap

The trap here is jumping to a remediation such as retraining or threshold tuning before confirming that the production inputs actually differ from the training distribution.

66
MCQeasy

A company wants to build a system that automatically tags uploaded images with objects they contain (e.g., 'car', 'tree', 'person'). Which AI application type is this?

A.Image classification/object detection
B.Recommendation system
C.Anomaly detection
D.Document intelligence
AnswerA

Object detection satisfies the requirement to tag multiple objects within one image, returning bounding boxes and labels per instance. Unlike image classification, which assigns a single label to the whole image, detection handles the stem's example of 'car', 'tree' and 'person' coexisting, so each object is identified separately.

Why this answer

The task of identifying and labeling objects (e.g., 'car', 'tree', 'person') within an image is a classic use case for image classification combined with object detection. Image classification assigns a single label to the entire image, while object detection localizes and classifies multiple objects within the image, which is exactly what the system requires.

Exam trap

The AI0-001 exam often tests the distinction between image classification (single label per image) and object detection (multiple localized objects), so candidates may mistakenly choose image classification alone when the question implies multiple objects per image.

How to eliminate wrong answers

Option B is wrong because recommendation systems analyze user behavior and preferences to suggest items (e.g., movies, products), not to identify objects in images. Option C is wrong because anomaly detection identifies unusual patterns or outliers in data (e.g., fraud detection), not the presence of common objects in images. Option D is wrong because document intelligence focuses on extracting text, structure, and information from documents (e.g., OCR, form processing), not on visual object recognition.

67
MCQeasy

A startup wants to add an AI-powered virtual assistant to their mobile app. They have limited in-house AI expertise and need a solution that can be integrated quickly with minimal infrastructure management. Which deployment pattern is MOST suitable?

A.Implement an asynchronous processing queue for all user requests
B.Train and deploy a custom model on an on-premises server
C.Deploy the model on edge devices for offline inference
D.Use a cloud-based AI microservice (e.g., Amazon Lex, Azure Bot Service) with a pre-built model
AnswerD

A cloud-based AI microservice with a pre-built model removes the need to train, host or scale models, directly satisfying the limited in-house expertise and minimal infrastructure management constraints. Integration is largely API configuration, enabling rapid deployment within the mobile app.

Why this answer

Using a cloud-based AI microservice with a pre-built model is most suitable because it requires minimal AI expertise and infrastructure management. The startup can integrate the service via APIs quickly, leveraging the provider's pre-trained models and scalability. This aligns with their need for rapid integration and limited in-house AI skills.

Exam trap

The trap is overcomplicating the solution by assuming custom model training is necessary; the exam often tests the ability to choose managed services when expertise and time are limited.

How to eliminate wrong answers

Option A is wrong because an asynchronous processing queue is an architectural pattern for decoupling components, not a deployment pattern for AI assistants, and it does not address the need for pre-built AI capabilities. Option B is wrong because training and deploying a custom model on-premises requires significant AI expertise and infrastructure management, contradicting the startup's constraints. Option C is wrong because deploying on edge devices for offline inference is complex, requires model optimization, and may not be necessary for a mobile app that can use cloud services.

68
MCQeasy

Which embedding type is MOST suitable for capturing semantic meaning of text in a RAG pipeline?

A.Bag-of-words vectors
B.Dense embeddings from a pre-trained transformer model
C.TF-IDF vectors
D.One-hot encoding
AnswerB

Pre-trained transformer embeddings encode contextual semantics into fixed dense vectors, so semantically similar passages land close together in vector space. This supports accurate similarity search and retrieval in a RAG pipeline, unlike sparse lexical vectors that capture term overlap rather than meaning.

Why this answer

Dense embeddings from pre-trained transformer models (e.g., BERT, Sentence-BERT) capture semantic meaning by representing text as high-dimensional vectors where similar meanings are close in vector space. This is essential for RAG pipelines to retrieve relevant documents based on semantic similarity rather than lexical overlap. Dense embeddings are the standard for modern semantic search.

Exam trap

The trap is confusing traditional sparse representations (TF-IDF, bag-of-words) with dense embeddings; candidates may think TF-IDF captures semantics because it weights terms, but it does not capture contextual meaning.

How to eliminate wrong answers

Option A is wrong because bag-of-words vectors ignore word order and semantics, relying only on word frequency, which fails to capture meaning. Option C is wrong because TF-IDF vectors are sparse and based on term frequency-inverse document frequency, which also ignores semantic relationships and context. Option D is wrong because one-hot encoding represents words as isolated categorical variables with no semantic similarity, making it unsuitable for semantic search.

69
Multi-Selecthard

A team is implementing a RAG system for a legal document Q&A. They need to chunk documents effectively. Which THREE chunking strategies should they consider to improve retrieval accuracy for legal texts that contain hierarchical sections (clauses, sub-clauses, definitions)?

Select 3 answers
A.Hierarchical chunking that indexes chunks at clause and sub-clause levels with parent relationships
B.Overlapping chunks with a 10% overlap between consecutive chunks
C.Fixed-size chunking with a 512-token window and no overlap
D.Chunking based on the document's table of contents and section hierarchy
E.Semantic chunking that splits at natural boundaries (e.g., section headings, paragraph breaks)
AnswersA, D, E

Indexing at clause and sub-clause levels with parent links preserves the legal hierarchy, so retrieval can return a precise sub-clause while still supplying its enclosing clause as context. This directly addresses the stem's hierarchical sections constraint.

Why this answer

Option A is correct because hierarchical chunking indexes chunks at clause and sub-clause levels while preserving parent relationships, which matches the legal document's nested structure and lets retrieval return a precise sub-clause while still offering the enclosing clause for context. Option D is correct because chunking based on the document's table of contents and section hierarchy aligns chunk boundaries with the document's own logical divisions, so each chunk corresponds to a meaningful legal section rather than an arbitrary token span. Option E is correct because semantic chunking splits at natural boundaries such as section headings and paragraph breaks, keeping definitions and clauses intact and avoiding mid-sentence or mid-clause cuts that would degrade embedding quality and retrieval accuracy.

Option B is not among the correct answers because a fixed 10% overlap is a generic heuristic that does not respect legal hierarchy and can duplicate or fragment clause text. Option C is not among the correct answers because fixed-size 512-token windows with no overlap ignore section structure and will frequently split clauses, sub-clauses, and definitions across chunk boundaries.

70
MCQmedium

An organization wants to implement an AI system to automatically categorize support tickets into predefined categories. They have a labeled dataset of 10,000 tickets. Which approach is MOST appropriate?

A.Use a rule-based system with keyword matching
B.Use a prompt-based LLM with few-shot examples
C.Fine-tune a pre-trained text classification model
D.Train a custom neural network from scratch
AnswerC

Fine-tuning a pre-trained text classification model leverages the 10,000 labelled tickets to adapt existing language representations to the organisation's specific categories, satisfying the stem's supervised categorisation requirement. This outperforms training from scratch on a small dataset and avoids the cost and latency of prompt-based large-model inference.

Why this answer

Fine-tuning a pre-trained text classification model is a standard and effective approach for supervised classification when labeled data is available.

71
MCQeasy

In the AI project lifecycle, which phase involves partitioning the dataset into training, validation, and test sets?

A.Data acquisition
B.Model selection
C.Data preparation
D.Problem definition
AnswerC

Data preparation covers cleaning, labelling and splitting the dataset into training, validation and test partitions before modelling begins. These subsets support fitting, hyperparameter tuning and unbiased final evaluation respectively, so the split belongs to this phase rather than data collection or deployment.

Why this answer

Data preparation includes splitting the data to evaluate model performance and prevent leakage.

72
MCQmedium

A company is building an AI-powered document intelligence system to extract key fields from scanned invoices. The data contains 95% of invoices from one vendor and 5% from others. During model training, the F1 score is 0.95 on the overall test set, but the performance on the minority vendor invoices is very poor. What is the MOST likely cause?

A.The model is overfitting on the minority class
B.The dataset is imbalanced, and the model is biased toward the majority class
C.The data has a train/test leakage problem
D.The feature extraction is incorrect for the minority vendor invoices
AnswerB

With 95% of invoices from one vendor, the model optimises for the majority class, so minority-vendor patterns are underweighted during training. High overall F1 masks this, since majority-class performance dominates the metric. The poor minority-vendor results stem directly from this class imbalance, not from overfitting or feature scaling.

Why this answer

The F1 score of 0.95 on the overall test set is misleading because 95% of invoices come from one vendor, so the model can achieve high overall accuracy by performing well on the majority class while failing on the minority class. This is a classic class imbalance problem where the model is biased toward the majority class. The poor performance on minority vendor invoices confirms that the model has not learned to generalize across vendors.

Exam trap

AI0-001 often tests the misconception that a high overall F1 score guarantees good performance across all classes, but candidates must recognize that imbalanced datasets can mask poor minority class performance.

How to eliminate wrong answers

Option A is wrong because overfitting on the minority class would typically cause poor performance on the majority class, not the other way around; here the model performs well on the majority and poorly on the minority, indicating bias toward the majority. Option C is wrong because train/test leakage would inflate performance on both majority and minority classes, not selectively harm the minority class. Option D is wrong because incorrect feature extraction for the minority vendor would likely affect all classes if the features are shared, and the symptom is specifically poor performance on the minority class, which is explained by imbalance.

73
MCQmedium

A retail bank is deploying a customer-facing AI assistant that must never disclose internal policy text. The team has a system prompt with instructions, but red-team testing shows users can extract the policy by asking the model to 'repeat everything above this line.' Which implementation change most directly mitigates this prompt-injection extraction risk in production?

A.Reduce the context window size so there is less room for the model to repeat the system prompt.
B.Move the policy text from the system prompt into the fine-tuning dataset so the model learns it as weights.
C.Increase the model temperature so responses become less deterministic and harder to reverse engineer.
D.Add an input guardrail that detects and blocks prompt-extraction patterns before the request reaches the model.
AnswerD

An input guardrail inspects the user prompt for known injection and extraction patterns, such as 'repeat everything above,' and blocks or sanitizes the request before it reaches the model. This directly addresses the attack vector shown in red-team testing while keeping the system prompt private, and it can be tuned and logged without retraining the underlying model.

Why this answer

The extraction attack succeeds because untrusted user input is concatenated with trusted instructions and the model follows the most recent instruction. Placing a guardrail in front of the model detects and blocks known extraction phrasing, which is the most direct runtime mitigation. Temperature, fine-tuning, and context size do not intercept the malicious instruction before it is processed.

Exam trap

The trap here is assuming that hiding or shortening the system prompt prevents extraction, when the real fix is validating untrusted input before it reaches the model.

74
MCQmedium

A data scientist is training a binary classifier and observes that the training accuracy is 99% but the test accuracy is only 70%. Which of the following is the MOST likely cause?

A.The model is underfitting the training data
B.The learning rate is too high
C.The model is overfitting the training data
D.The test set contains data leakage from the training set
AnswerC

A 29-point gap between training and test accuracy is the classic signature of overfitting: the model has memorised training noise rather than learning generalisable patterns. The constraint in the stem is the large train-test performance divergence, which variance-reduction techniques such as regularisation or more data would address.

Why this answer

A large gap between training accuracy (99%) and test accuracy (70%) is the classic signature of overfitting: the model has memorized the training data, including its noise, and therefore fails to generalize to unseen examples. The high training score confirms the model has enough capacity to fit the training set, while the poor test score shows that capacity is being used for memorization rather than learning generalizable patterns.

Exam trap

AI0-001 often tests the train-vs-test accuracy gap pattern, and candidates confuse overfitting (high train, low test) with underfitting (low train, low test) or with data leakage (which inflates test scores).

How to eliminate wrong answers

Option A is wrong because underfitting produces low accuracy on BOTH training and test sets, not a 99% training score. Option B is wrong because an excessively high learning rate typically causes unstable or diverging loss curves, not a clean 99%/70% train-test split. Option D is wrong because data leakage from the test set into training would inflate test accuracy (making it suspiciously high), not depress it to 70%.

75
Multi-Selectmedium

A financial services firm is deploying an LLM-based assistant that summarizes internal earnings reports. Compliance requires that the assistant's outputs be auditable and that sensitive financial figures not be sent to an external model provider. Which TWO implementation measures should the team adopt? (Choose two.)

Select 2 answers
A.Deploy the model in a private environment where inference occurs within the firm's network boundary.
B.Enable the model provider's zero-retention API option so prompts are not stored by the vendor.
C.Log each prompt, retrieved context, model version, and response with timestamps for later review.
D.Restrict the assistant to answering only yes-or-no questions about the earnings reports.
E.Increase the model's temperature to diversify summaries so no single output is treated as authoritative.
AnswersA, C

Running inference inside the firm's network prevents sensitive financial figures from leaving the organization, directly satisfying the data-residency requirement. It also gives the firm control over model versions and logging infrastructure, which supports auditability. External API calls would transmit the same sensitive content to a third party.

Why this answer

Auditability requires durable records of prompts, context, model versions, and responses, while data-residency requires that sensitive figures never leave the firm's boundary. Running inference in a private environment and logging every interaction together satisfy both requirements. The remaining options either still transmit data externally or do nothing for auditability.

Exam trap

The trap here is accepting a vendor-side zero-retention promise as equivalent to keeping data internal, when the figures are still transmitted to an external provider.

Page 1 of 2 · 132 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Aio Implementing Ai questions.