Courseiva

CCNA Ai Models Data Engineering Questions

67 questions · Ai Models Data Engineering topic · All types, answers revealed

1
MCQhard

A data scientist trains a deep learning model on a large dataset. The training loss decreases steadily but the validation loss starts increasing after 20 epochs. The scientist uses early stopping with patience=5. Which of the following is the MOST likely cause and best corrective action?

A.Model is overfitting; add dropout regularization.
B.Training data is not representative; collect more data.
C.Model is underfitting; increase model capacity.
D.Learning rate too high; reduce learning rate.
AnswerA

Overfitting is the cause: training loss falls while validation loss rises, so the model memorises training data rather than generalising. Dropout randomly deactivates neurons during training, constraining the network's capacity and reducing this divergence, satisfying the stem's requirement to address the validation-loss increase directly.

Why this answer

The training loss decreasing while validation loss increasing after 20 epochs is a classic sign of overfitting, where the model memorizes training data noise instead of generalizing. Early stopping with patience=5 would halt training after 5 epochs of no validation improvement, but the root cause is overfitting. Adding dropout regularization randomly drops neurons during training, forcing the network to learn more robust features and reducing overfitting.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by showing a diverging validation loss curve, and the trap here is that candidates may confuse overfitting with a learning rate issue or data quality problem, leading them to choose 'reduce learning rate' or 'collect more data' instead of the correct regularization technique.

How to eliminate wrong answers

Option B is wrong because the validation loss increasing while training loss decreases indicates overfitting, not unrepresentative data; collecting more data might help but is not the most direct corrective action for overfitting. Option C is wrong because underfitting would show high training loss that does not decrease, not a decreasing training loss with increasing validation loss. Option D is wrong because a high learning rate would typically cause training loss to oscillate or diverge, not steadily decrease; reducing learning rate addresses convergence issues, not overfitting.

2
MCQeasy

A logistics company uses a machine learning model to predict delivery times based on historical data. The model was performing well, but recently it started making inaccurate predictions, especially for routes that have experienced new traffic patterns and road closures. The data engineering team receives an alert that the model's accuracy has dropped by 15% over the last week. They suspect data drift. The team has access to the original training data and a continuous stream of new data. What is the most appropriate first step for the team to take?

A.Roll back the model to the previous stable version and schedule a full audit of the data pipeline.
B.Compare the distributions of key features between the training data and the recent data to quantify data drift.
C.Immediately retrain the model using the most recent data to adapt to the new patterns.
D.Add more features to the model to capture the new traffic patterns and road closures.
AnswerB

Distribution comparison directly quantifies drift by contrasting feature statistics between the original training data and recent streamed data, confirming whether new traffic patterns and closures shifted inputs. This diagnostic precedes retraining, isolating whether the 15% accuracy drop stems from input drift rather than label or concept change.

Why this answer

The first step in diagnosing a suspected data drift is to statistically compare the distributions of key features between the training data and the recent streaming data. This quantifies whether the input data distribution has changed, which directly explains the accuracy drop. Without this analysis, any corrective action (like retraining or rollback) would be premature and could mask the root cause.

Exam trap

CompTIA often tests the misconception that the immediate response to a performance drop should be retraining or rollback, rather than first diagnosing the type of drift (data drift vs. concept drift) through distribution comparison.

How to eliminate wrong answers

Option A is wrong because rolling back the model without first confirming data drift wastes time and may not address the new traffic patterns; it assumes the previous model is still valid, which is false if drift is present. Option C is wrong because immediately retraining on recent data without verifying drift could introduce bias or overfit to transient noise, and it ignores the need to first understand what changed. Option D is wrong because adding features without first analyzing drift is a blind attempt that may not solve the distribution shift and could increase model complexity unnecessarily.

3
MCQmedium

A machine learning team is deploying a sentiment analysis model for customer reviews. The model was trained on reviews from an e-commerce site but will be used for a social media platform. The team observes a drop in accuracy. Which concept best explains this issue?

A.Data drift
B.Concept drift
C.Bias-variance tradeoff
D.Overfitting
AnswerA

The model encounters social media text whose vocabulary, length and style differ from the e-commerce reviews it was trained on, so the input distribution shifts between training and deployment. This covariate shift is data drift, explaining the accuracy drop without any change in the underlying sentiment-label relationship.

Why this answer

Data drift occurs when the statistical properties of the input data change between the training and production environments. Here, the model was trained on e-commerce reviews but is now processing social media posts, which have different vocabulary, tone, and structure, causing a mismatch in the input distribution and leading to accuracy degradation.

Exam trap

CompTIA often tests the distinction between data drift (input distribution change) and concept drift (relationship change), and candidates mistakenly choose concept drift when the scenario describes a change in the input data source rather than a change in the underlying mapping from inputs to outputs.

How to eliminate wrong answers

Option B is wrong because concept drift refers to a change in the underlying relationship between input features and the target variable over time, not a change in the input data distribution itself. Option C is wrong because bias-variance tradeoff is a model selection concept describing the balance between underfitting and overfitting, not an explanation for performance drop due to data distribution shift. Option D is wrong because overfitting occurs when a model learns training data too well, including noise, and fails to generalize to new data from the same distribution, not to a different distribution.

4
MCQmedium

A retail analytics team is preparing a dataset of product reviews for a sentiment classification model. The dataset contains 50,000 reviews, but only 2,000 are labeled as positive or negative. The team wants to use the unlabeled reviews to improve model performance. Which approach best leverages the unlabeled data?

A.Apply principal component analysis (PCA) to reduce the dimensionality of the unlabeled reviews and then train a classifier on the reduced features.
B.Apply semi-supervised learning using a self-training algorithm that iteratively labels high-confidence unlabeled examples and retrains the model.
C.Perform data augmentation by generating synthetic reviews using a generative adversarial network (GAN) trained on the labeled data.
D.Use transfer learning by fine-tuning a pre-trained language model on the 2,000 labeled reviews only, ignoring the unlabeled data.
AnswerB

Self-training is a semi-supervised technique that uses a small labeled set to train an initial model, then predicts labels for unlabeled data, adds the most confident predictions to the training set, and repeats. This directly uses the 48,000 unlabeled reviews to improve the model without manual labeling, making it the best fit for the scenario.

Why this answer

Semi-supervised learning, specifically self-training, is designed to exploit a small labeled set alongside a large unlabeled set. By iteratively adding high-confidence pseudo-labels from the unlabeled reviews, the model can learn from the broader data distribution, improving generalization. The other options either ignore the unlabeled data or use techniques that are not appropriate for text sentiment classification with limited labels.

Exam trap

The trap here is assuming that any technique using unlabeled data, such as PCA, is sufficient, when in fact only semi-supervised methods directly incorporate unlabeled data into the training process to improve classification.

5
Multi-Selectmedium

Which THREE practices are recommended for versioning machine learning models in a production environment?

Select 3 answers
A.Use a model registry like MLflow or DVC.
B.Store model metadata such as hyperparameters and training data hash.
C.Automate model deployment based on version tags.
D.Use Git to version model binaries.
E.Keep only the latest model to save storage.
AnswersA, B, C

A model registry such as MLflow or DVC provides centralised, immutable version tracking, linking each model artefact to its training data, parameters and metrics. This satisfies the production requirement for reproducible lineage and rollback, letting teams promote or revert specific model versions without ambiguity.

Why this answer

Option A is correct because a dedicated model registry such as MLflow or DVC is purpose-built for tracking model artifacts, versions, and lifecycle stages, which is the recommended practice for production ML versioning. Option B is correct because storing metadata like hyperparameters and the training data hash ties each model version to its exact training configuration and dataset, enabling reproducibility and auditability. Option C is correct because automating deployment based on version tags ensures that only approved, traceable model versions reach production and keeps deployment consistent with the registry.

Option D is not recommended because Git is designed for source code and text, not large binary model files, which bloat repositories and lack proper artifact lineage. Option E is wrong because discarding older models destroys rollback capability, reproducibility, and compliance auditing, and storage savings do not justify that risk.

Exam trap

CompTIA often tests the misconception that Git is suitable for versioning all artifacts, including large binary model files, when in fact Git's architecture is optimized for text diffs and cannot efficiently manage model binaries in a production ML pipeline.

6
MCQhard

A team is training a deep learning model for image classification. The training loss decreases rapidly but validation loss starts increasing after a few epochs. Which regularization technique should be applied to mitigate this issue?

A.Data augmentation
B.L2 regularization
C.Early stopping
D.Dropout
AnswerC

Rising validation loss alongside falling training loss signals overfitting. Early stopping halts training at the epoch where validation loss is minimal, restoring the best generalising weights. This directly mitigates the divergence described, unlike dropout or weight decay, which alter the architecture or loss function.

Why this answer

Early stopping halts training when validation loss starts increasing, preventing overfitting. Option A (data augmentation) is wrong because it increases data diversity but does not stop training when validation loss increases. Option B (L2 regularization) is wrong because it penalizes large weights but does not directly address the issue of validation loss increasing.

Option D (dropout) is wrong because while it helps generalize by randomly dropping neurons, it does not stop training when overfitting occurs.

7
MCQmedium

A machine learning engineer is building a model to predict whether a customer will make a purchase within the next week. The dataset contains 10,000 samples with 20 features, and the target variable is binary. The engineer wants to use a model that provides interpretable results to explain predictions to business stakeholders. Which model is most appropriate?

A.Gradient boosting machine
B.Logistic regression
C.Random forest
D.Support vector machine with RBF kernel
AnswerB

Logistic regression is a linear model that provides coefficients for each feature, indicating the direction and magnitude of their influence on the predicted probability. This makes it highly interpretable, as stakeholders can understand how each feature contributes to the prediction. It is well-suited for binary classification tasks and performs reasonably well when the relationship between features and the log-odds is approximately linear. In this scenario, with 20 features and 10,000 samples, logistic regression can be trained efficiently and the coefficients can be explained to business users.

Why this answer

Logistic regression is a linear model that provides coefficients for each feature, making it highly interpretable. Business stakeholders can understand how each feature influences the predicted probability of purchase. While ensemble methods like random forest and gradient boosting may offer higher accuracy, they are less transparent.

Support vector machines with non-linear kernels are also difficult to interpret. Therefore, logistic regression is the most appropriate choice when interpretability is a primary requirement.

Exam trap

The trap here is assuming that more complex models always yield better business value, overlooking the need for interpretability in stakeholder communication.

8
MCQeasy

A data scientist notices that a binary classification model consistently predicts the majority class. Which data engineering technique should be applied?

A.Feature scaling
B.Dimensionality reduction
C.Polynomial features
D.Oversampling
AnswerD

Consistent majority-class prediction indicates severe class imbalance, where the loss is minimised by always guessing the dominant label. Oversampling replicates or synthesises minority-class examples, rebalancing the training distribution so the classifier learns the minority decision boundary.

Why this answer

Oversampling (Option D) is correct because the model's bias toward the majority class indicates a class imbalance problem. By synthetically increasing the number of minority class samples (e.g., using SMOTE or random oversampling), the training data becomes more balanced, allowing the classifier to learn decision boundaries that are not skewed toward the majority class.

Exam trap

CompTIA often tests the misconception that feature scaling or dimensionality reduction can fix class imbalance, when in reality these techniques address different issues like feature magnitude or curse of dimensionality, not skewed target distributions.

How to eliminate wrong answers

Option A is wrong because feature scaling normalizes the range of input features (e.g., via min-max scaling or standardization) but does not address class imbalance; it only prevents features with larger magnitudes from dominating gradient-based optimization. Option B is wrong because dimensionality reduction (e.g., PCA or t-SNE) reduces the number of features to combat overfitting or noise, but it does not alter the class distribution, so the majority class bias remains. Option C is wrong because polynomial features create interaction or higher-degree terms from existing features to capture non-linear relationships, but they do not change the ratio of majority to minority samples, leaving the imbalance untouched.

9
MCQmedium

A data pipeline processes customer data from multiple sources. The data quality check reveals duplicate records. Which step should the pipeline include to handle this?

A.Data deduplication
B.Data encryption
C.Data transformation
D.Data validation
AnswerA

Data deduplication removes duplicate records by identifying matching rows across sources and retaining a single canonical entry, directly resolving the duplicate records the quality check flagged. It satisfies the stem's constraint of handling duplicates within the pipeline, unlike profiling or validation, which detect issues but do not eliminate them.

Why this answer

Duplicate records in a data pipeline compromise data integrity and downstream analytics. Data deduplication (Option A) is the correct step because it identifies and removes redundant entries based on key fields or fuzzy matching, ensuring each customer record is unique. This is a core data quality operation in ETL pipelines, often implemented via hash-based comparison or SQL window functions like ROW_NUMBER().

Exam trap

This question tests the distinction between data quality actions (deduplication) and data security or formatting actions (encryption, transformation), leading candidates to confuse validation (which only flags issues) with remediation (which removes duplicates).

How to eliminate wrong answers

Option B (Data encryption) is wrong because encryption secures data at rest or in transit (e.g., AES-256, TLS 1.3) and does not address duplicate records. Option C (Data transformation) is wrong because transformation changes data format or structure (e.g., type casting, normalization) but does not inherently remove duplicates. Option D (Data validation) is wrong because validation checks data against rules (e.g., schema constraints, range checks) and flags errors, but it does not actively eliminate duplicate rows.

10
Multi-Selecteasy

A data engineer is preparing a dataset for a binary classification model. The dataset has 10,000 samples with 100 features. To improve model performance and reduce training time, the engineer decides to perform feature selection. Which two techniques are appropriate for this task? (Select TWO).

Select 2 answers
A.Normalization
B.Recursive Feature Elimination (RFE)
C.L1 Regularization
D.One-Hot Encoding
E.Principal Component Analysis (PCA)
AnswersB, C

Recursive Feature Elimination fits a model, ranks features by importance, removes the weakest, and repeats, progressively shrinking the 100-feature set. This reduces training time and can improve performance by eliminating irrelevant or noisy predictors from the dataset.

Why this answer

Recursive Feature Elimination (RFE) (B) is a wrapper-based feature selection method that repeatedly trains a model, ranks features by importance (e.g., coefficients or feature importances), and prunes the least important ones until the desired number of features remains, directly reducing the 100-feature space to improve performance and cut training time. L1 Regularization (C) — Lasso — adds a penalty equal to the absolute value of coefficients to the loss function, which drives many feature coefficients exactly to zero, effectively performing embedded feature selection and yielding a sparse model. Normalization (A) is a scaling preprocessing step (e.g., min-max or z-score) that changes feature magnitudes but does not remove features, so it is not feature selection.

One-Hot Encoding (D) is a categorical-encoding transformation that expands categorical variables into binary columns, increasing dimensionality rather than reducing it. Principal Component Analysis (E) is a dimensionality-reduction technique that creates new uncorrelated components from linear combinations of the original features, but it is not feature selection because it does not retain or select the original features.

Exam trap

CompTIA often tests the distinction between feature selection (keeping original features) and dimensionality reduction (creating new features), so candidates mistakenly select PCA thinking it selects features, when it actually transforms them into principal components.

11
MCQmedium

A data engineer needs to design a data pipeline for a real-time fraud detection system. The system requires low-latency processing of streaming transactions. Which architecture is most appropriate?

A.Stream processing with Apache Kafka and Flink
B.Data lake with Apache Spark
C.Batch processing with Apache Hadoop
D.Microservices architecture with REST APIs
AnswerA

Kafka ingests high-throughput transaction streams durably, while Flink performs stateful, event-time windowed computation with millisecond latency, satisfying the low-latency streaming requirement. Batch architectures such as Hadoop or Spark microbatching introduce latency unsuited to real-time fraud detection.

Why this answer

Apache Kafka provides a distributed, fault-tolerant event streaming platform that ingests high-throughput transaction data with low latency, while Apache Flink offers true stream processing with exactly-once semantics and sub-second event-time processing. Together, they enable real-time fraud detection by analyzing transactions as they arrive, without the delays inherent in batch or micro-batch approaches.

Exam trap

CompTIA often tests the distinction between true stream processing (e.g., Flink, Kafka Streams) and micro-batch or near-real-time processing (e.g., Spark Streaming), where candidates mistakenly assume that any 'streaming' API (like Spark Streaming) is equivalent to low-latency stream processing.

How to eliminate wrong answers

Option B is wrong because a data lake with Apache Spark typically relies on micro-batch processing (e.g., Spark Streaming with a minimum batch interval of ~100ms), which introduces higher latency than true stream processing and is unsuitable for sub-second fraud detection. Option C is wrong because batch processing with Apache Hadoop (e.g., MapReduce) is designed for high-throughput, high-latency processing of large static datasets, not for real-time streaming where transactions must be evaluated within milliseconds. Option D is wrong because microservices architecture with REST APIs is a design pattern for building distributed services, not a data pipeline technology; REST APIs introduce synchronous request-response overhead and cannot natively handle continuous, unbounded data streams with low-latency stateful processing.

12
MCQhard

A machine learning engineer is preparing a dataset for a natural language processing task. The dataset contains text reviews with varying lengths, and the engineer plans to use a transformer model. Which preprocessing step is most critical to ensure the model can handle the input effectively?

A.Convert all text to lowercase and remove punctuation to reduce vocabulary size.
B.Perform stemming or lemmatization to reduce words to their base forms.
C.Tokenize the text into subword units and pad or truncate sequences to a fixed maximum length.
D.Apply one-hot encoding to each word in the vocabulary to create binary vectors.
AnswerC

Transformer models require input sequences of uniform length within a batch. Tokenization into subword units (e.g., WordPiece or BPE) handles out-of-vocabulary words and reduces vocabulary size. Padding shorter sequences and truncating longer ones to a fixed maximum length ensures batch processing. This is essential for efficient training and inference with transformers.

Why this answer

Transformer models require fixed-length input sequences, so tokenizing into subword units and padding/truncating to a uniform length is essential. This enables batch processing and handles out-of-vocabulary words. Other steps like lowercasing, one-hot encoding, or stemming are either not critical or incompatible with transformer input expectations.

Exam trap

The trap here is focusing on traditional text normalization steps like lowercasing or stemming while overlooking the transformer-specific need for uniform sequence lengths via tokenization and padding.

13
MCQhard

A fraud detection model has high precision but low recall. The cost of false negatives is very high. Which threshold adjustment should be made?

A.Use class weights during training
B.Apply SMOTE to the training data
C.Decrease classification threshold
D.Increase classification threshold
AnswerC

Lowering the threshold classifies more cases as fraudulent, raising recall and cutting costly false negatives. Precision falls as a result, but the stem states false negatives carry very high cost, so this trade-off is appropriate.

Why this answer

Decreasing the classification threshold makes the model more sensitive, classifying more instances as positive. This increases recall by catching more true positives, directly addressing the high cost of false negatives, even though precision may drop.

Exam trap

The AI0-001 exam often tests the distinction between training-time techniques (like class weights or SMOTE) and post-training threshold tuning, trapping candidates who confuse data-level remedies with decision boundary adjustments.

How to eliminate wrong answers

Option A is wrong because using class weights during training rebalances the loss function to penalize false negatives more, which is a training-time adjustment, not a post-training threshold change. Option B is wrong because SMOTE oversamples the minority class in the training data to address class imbalance, which is a data preprocessing step, not a threshold adjustment. Option D is wrong because increasing the classification threshold makes the model more conservative, reducing false positives but further lowering recall, which worsens the false negative problem.

14
MCQhard

A machine learning team is developing a model to predict server failure from telemetry data. They use a deep neural network with 3 hidden layers. After training, the model achieves 99% accuracy on training data but only 85% on validation data. Which technique should the team apply to reduce the generalization error?

A.Increase the number of hidden layers
B.Apply L2 regularization
C.Increase the learning rate
D.Add more training data
AnswerB

L2 regularization adds a penalty on large weights to the loss function, shrinking model complexity and curbing the overfitting behind the 99% training versus 85% validation gap. This directly reduces the generalization error the team needs to lower.

Why this answer

The model exhibits high variance (overfitting) because it achieves 99% accuracy on training data but only 85% on validation data. L2 regularization (also known as weight decay) adds a penalty proportional to the squared magnitude of the weights to the loss function, which discourages the network from fitting noise in the training data and improves generalization. This directly reduces the gap between training and validation performance.

Exam trap

CompTIA often tests the distinction between techniques that address overfitting (regularization) versus those that address underfitting (more layers, higher learning rate) or data quantity, leading candidates to mistakenly choose adding more data or increasing model complexity.

How to eliminate wrong answers

Option A is wrong because increasing the number of hidden layers would increase model capacity, making overfitting worse and further increasing generalization error. Option C is wrong because increasing the learning rate can cause the optimizer to overshoot minima or diverge, but it does not directly address overfitting; it may even prevent convergence. Option D is wrong because while adding more training data can help reduce overfitting, it is not the most direct or practical technique when the team already has a model that overfits; regularization is a more immediate and targeted solution.

15
MCQeasy

A data scientist is working on a project to classify images of handwritten digits. The dataset consists of 60,000 training images and 10,000 test images, each 28x28 pixels in grayscale. The scientist wants to build a model that can automatically extract features and achieve high accuracy. Which type of model is most suitable for this task?

A.K-means clustering
B.Decision tree
C.Convolutional neural network (CNN)
D.Logistic regression
AnswerC

Convolutional neural networks are specifically designed for image data. They use convolutional layers to automatically learn spatial hierarchies of features, such as edges, textures, and shapes, from raw pixel values. This makes them highly effective for image classification tasks like handwritten digit recognition. CNNs also benefit from parameter sharing and local connectivity, reducing the number of parameters compared to fully connected networks. Given the image size and dataset, a CNN can achieve high accuracy with reasonable computational resources. Therefore, a CNN is the most suitable model.

Why this answer

Convolutional neural networks are designed for image data and can automatically learn relevant features through convolutional layers. They are highly effective for handwritten digit classification. Logistic regression requires manual feature engineering, decision trees struggle with high-dimensional image data, and K-means is unsupervised and not suitable for classification.

Therefore, a CNN is the most suitable model for this task.

Exam trap

The trap here is selecting a simpler model like logistic regression due to familiarity, without recognizing the need for automatic feature extraction in image data.

16
MCQeasy

A data scientist is preparing a dataset for a machine learning model and notices that one feature has a range from 0 to 1,000,000, while another feature ranges from 0 to 1. The model to be used is a k-nearest neighbors (KNN) classifier. Which preprocessing step is MOST important to apply before training?

A.Principal component analysis (PCA) to reduce dimensionality.
B.One-hot encoding for all features to convert them into binary vectors.
C.Removing outliers from the feature with the larger range.
D.Feature scaling, such as min-max normalization or standardization.
AnswerD

KNN relies on distance calculations between data points. If one feature has a much larger range than others, it will dominate the distance metric, making the model effectively ignore the smaller-range features. Scaling ensures all features contribute equally to distance computations, which is essential for KNN to perform well.

Why this answer

KNN uses distance metrics like Euclidean distance, which are sensitive to feature scales. A feature ranging from 0 to 1,000,000 will have a much larger impact on distance than one ranging from 0 to 1, effectively drowning out the smaller feature. Applying feature scaling, such as min-max normalization or standardization, ensures that all features contribute proportionally to the distance, leading to a more accurate and balanced model.

Exam trap

The trap here is assuming that dimensionality reduction or outlier removal is the primary fix, when the core issue is the disparity in feature scales that directly affects distance-based algorithms like KNN.

17
MCQeasy

A company streams sensor data from IoT devices. The data arrives as JSON messages at high velocity. Which data pipeline architecture is BEST suited to handle this streaming data for near-real-time analytics?

A.Batch processing using Hadoop MapReduce every 24 hours.
B.Batch processing using nightly ETL jobs.
C.Single-node database with periodic inserts.
D.Stream processing using Apache Kafka and Spark Streaming.
AnswerD

Apache Kafka ingests the high-velocity JSON sensor messages durably as a distributed log, while Spark Streaming consumes those partitions and performs micro-batch analytics, satisfying the near-real-time requirement. Unlike batch pipelines, this architecture processes each message as it arrives rather than waiting for scheduled windows.

Why this answer

Apache Kafka acts as a distributed, fault-tolerant ingestion layer that can handle high-velocity JSON messages, while Spark Streaming processes the data in micro-batches for near-real-time analytics. This combination provides the low-latency, scalable pipeline required for streaming IoT sensor data, unlike batch or single-node approaches.

Exam trap

CompTIA often tests the distinction between batch and stream processing by presenting batch options that seem 'reliable' or 'traditional,' trapping candidates who overlook the explicit 'near-real-time' requirement in the question.

How to eliminate wrong answers

Option A is wrong because Hadoop MapReduce is designed for batch processing of large static datasets, not for continuous high-velocity streaming data, and a 24-hour cycle cannot meet near-real-time requirements. Option B is wrong because nightly ETL jobs introduce hours of latency, making them unsuitable for near-real-time analytics on streaming data. Option C is wrong because a single-node database with periodic inserts cannot scale to handle high-velocity IoT data streams and will become a bottleneck, failing to provide near-real-time processing.

18
MCQmedium

An organization uses a machine learning model to approve loans. The model shows higher false positive rates for a protected group. Which data engineering step should be taken to mitigate this?

A.Remove the protected attribute from training data
B.Use adversarial debiasing technique
C.Increase model complexity
D.Add synthetic data to balance groups
AnswerB

Adversarial debiasing trains the model alongside an adversary that predicts the protected attribute, forcing learned representations to be independent of it. This reduces the disparate false positive rates, directly mitigating the bias the organisation observed against the protected group.

Why this answer

Adversarial debiasing is a technique that trains the model to minimize prediction error while simultaneously preventing an adversary from predicting the protected attribute from the model's outputs. This directly reduces disparate impact by forcing the model to learn representations that are uncorrelated with the protected group, thereby lowering false positive rates for that group without simply removing the attribute.

Exam trap

A common misconception tested in CompTIA AI is that removing the protected attribute is sufficient to eliminate bias, when in reality proxy features and correlated variables can perpetuate discrimination, making adversarial debiasing a more robust solution.

How to eliminate wrong answers

Option A is wrong because simply removing the protected attribute from training data does not eliminate proxy features (e.g., zip code, income) that correlate with the protected group, so bias can persist through correlated features. Option C is wrong because increasing model complexity typically exacerbates overfitting and can amplify existing biases rather than mitigate them, as the model may learn spurious correlations tied to the protected group. Option D is wrong because adding synthetic data to balance groups addresses class imbalance but does not directly correct the model's decision boundary bias that causes higher false positives for a specific group; it may even introduce artifacts if synthetic data is not carefully generated.

19
MCQhard

An e-commerce company needs to update its recommendation model continuously as user preferences change. The model currently retrains from scratch every night, but the training time is too long. Which approach would reduce training time while keeping the model up-to-date?

A.Use dimensionality reduction on features.
B.Implement incremental learning using online gradient descent.
C.Switch to a simpler model.
D.Increase the batch size for retraining.
AnswerB

Online gradient descent updates weights from each new sample or mini-batch, so the model adapts without a full retraining pass. This directly cuts the nightly training time constraint while keeping recommendations current as user preferences drift.

Why this answer

Incremental learning using online gradient descent updates the model parameters with each new data point or mini-batch, avoiding the need to retrain from scratch. This approach significantly reduces training time while continuously adapting to changing user preferences, making it ideal for real-time recommendation systems.

Exam trap

CompTIA often tests the misconception that dimensionality reduction or simpler models are the primary solution for reducing training time, when in fact incremental learning directly addresses the need for continuous updates without full retraining.

How to eliminate wrong answers

Option A is wrong because dimensionality reduction reduces the number of features but does not eliminate the need to retrain the entire model from scratch each night; the training time savings are marginal and the core problem of full retraining remains. Option C is wrong because switching to a simpler model may reduce training time but typically sacrifices model accuracy and expressiveness, which is critical for capturing nuanced user preferences in recommendations. Option D is wrong because increasing the batch size for retraining can actually increase memory usage and may not reduce overall training time if the model still retrains from scratch nightly; it does not address the fundamental inefficiency of full retraining.

20
MCQeasy

A data scientist is preparing a dataset for a natural language processing task. The dataset contains a 'review_text' column with free-form customer reviews. Before feeding the text into a machine learning model, the team wants to convert the text into numerical features. Which technique is most appropriate for this purpose?

A.Use one-hot encoding on each unique word in the 'review_text' column.
B.Convert each review to its average word length and use that as a single numerical feature.
C.Apply TF-IDF (Term Frequency-Inverse Document Frequency) vectorization to the 'review_text' column.
D.Hash each review to a fixed-length integer using a cryptographic hash function.
AnswerC

TF-IDF converts text into numerical vectors by weighting terms based on their frequency in a document and their rarity across the corpus. This highlights important words while downweighting common words like 'the' or 'and'. It is a standard and effective method for transforming unstructured text into features suitable for many machine learning algorithms, especially when the dataset is not extremely large.

Why this answer

TF-IDF is the most appropriate because it transforms text into numerical vectors that reflect term importance, capturing semantic content while reducing the influence of common words. It is widely used for text classification and clustering, and it works well with many machine learning algorithms without requiring deep learning architectures.

Exam trap

The trap here is thinking that any numerical conversion works, but techniques like hashing or average word length discard the semantic content needed for NLP tasks.

21
MCQhard

A medical imaging team is developing an AI model to detect tumors from CT scans. They have 10,000 labeled scans, but the labels were created by a semi-automated process with an estimated 20% error rate (mislabeled tumor vs. no tumor). The team trains a convolutional neural network (CNN) and achieves 90% accuracy on a held-out test set that was carefully validated by an expert radiologist. However, when deployed to a new hospital's patient population, the accuracy drops to 70%. The team suspects domain shift and label noise. Which strategy is most likely to improve model robustness for the new hospital?

A.Use active learning to select the most uncertain predictions from the new hospital's data, then have an expert radiologist correct those labels
B.Randomly select 1,000 scans from the new hospital and have them re-labeled by the radiologist
C.Collect 20,000 more scans with the same semi-automated labeling process
D.Reduce the CNN's number of layers and apply dropout to combat overfitting
AnswerA

Active learning targets the new hospital's own distribution, and expert correction removes label noise in exactly those uncertain cases. This adapts the model to the shifted domain while fixing the 20% labelling error, addressing both suspected causes.

Why this answer

Active learning selects the most uncertain predictions from the new hospital's data, allowing an expert radiologist to efficiently correct the most informative labels. This directly addresses both label noise (by correcting mislabeled examples) and domain shift (by focusing on samples where the model is uncertain in the new domain). Option B is wrong because random selection may not target the most impactful errors, wasting expert effort.

Option C is wrong because adding more noisy labels from the same flawed process will amplify label noise without correcting the domain shift. Option D is wrong because reducing model complexity and dropout are regularization techniques that do not fix label noise or domain shift.

22
MCQhard

A machine learning team is deploying a model that predicts loan default probabilities. The model outputs a probability score, and the team wants to convert it into a binary decision (default/no default). The costs of false positives and false negatives are not equal; a false negative (predicting no default when the customer defaults) is five times more costly than a false positive. Which approach best optimizes the decision threshold?

A.Set the threshold to 0.5 to balance false positives and false negatives equally.
B.Use the threshold that maximizes overall accuracy on the validation set.
C.Choose a threshold that minimizes the expected cost, weighting false negatives five times more than false positives.
D.Set the threshold to the prevalence of defaults in the training data.
AnswerC

The optimal threshold minimizes expected cost given the cost matrix. By weighting false negatives five times more, the threshold will be lowered to predict more defaults, reducing costly false negatives. This approach directly incorporates business costs and yields the most cost-effective decisions.

Why this answer

The optimal decision threshold minimizes expected cost, which requires weighting false negatives according to their higher cost. Lowering the threshold increases the number of predicted defaults, reducing expensive false negatives at the expense of more false positives. This cost-sensitive approach aligns model decisions with business objectives, unlike accuracy maximization or arbitrary thresholds.

Exam trap

The trap here is defaulting to 0.5 or accuracy-based thresholds, ignoring that the cost of a false negative is five times that of a false positive.

23
MCQmedium

A data engineer is designing a pipeline to ingest high-velocity clickstream events from a web application into a data lake. The events must be queryable within minutes of arrival, and the schema evolves frequently as new fields are added. Which storage approach best meets these requirements?

A.Batch load events every 24 hours into a data warehouse using a rigid ETL process with predefined transformations.
B.Store raw JSON files in an object storage bucket partitioned by date, and query them directly using a serverless SQL engine.
C.Load events into a relational database with a fixed schema, using ALTER TABLE to add columns as new fields appear.
D.Write events to a distributed message queue and retain them for 7 days, querying the queue directly for analytics.
AnswerB

This approach supports schema-on-read, allowing new fields to appear without rewriting existing data. Partitioning by date improves query performance and reduces scanned data. Serverless SQL engines can query JSON directly, enabling near-real-time analysis within minutes. It is cost-effective and scales with event volume, making it ideal for evolving clickstream data.

Why this answer

Storing raw JSON in object storage and querying with a serverless SQL engine provides schema-on-read flexibility and near-real-time access. It accommodates evolving schemas without costly migrations and scales to high volumes. The other options impose rigid schemas, high latency, or lack analytical query capabilities, making them unsuitable for this scenario.

Exam trap

The trap here is assuming that a relational database or message queue can serve as a scalable, queryable store for evolving, high-velocity data without significant drawbacks.

24
MCQeasy

During feature engineering, a data scientist creates a new feature that is a linear combination of two existing features. What risk does this pose to the model?

A.Multicollinearity
B.Data leakage
C.Overfitting
D.Underfitting
AnswerA

A feature built as a linear combination of two existing features is perfectly correlated with them, producing multicollinearity. This inflates the variance of coefficient estimates, making them unstable and hard to interpret, which is the specific risk the engineered linear combination introduces.

Why this answer

Creating a new feature as a linear combination of two existing features introduces perfect multicollinearity, where the new feature is an exact linear function of the original ones. This violates the assumption of no perfect multicollinearity in linear models, causing the design matrix to become singular and making coefficient estimates unstable or impossible to compute. Even in non-linear models, high multicollinearity can inflate variance and reduce interpretability.

Exam trap

CompTIA often tests the distinction between multicollinearity and overfitting, trapping candidates who confuse feature redundancy with model complexity.

How to eliminate wrong answers

Option B is wrong because data leakage refers to using information from outside the training set (e.g., future data or target leakage), not to relationships among features within the training data. Option C is wrong because overfitting is caused by a model learning noise or overly complex patterns, not by linear dependencies between features; multicollinearity primarily affects coefficient stability, not generalization error directly. Option D is wrong because underfitting occurs when a model is too simple to capture underlying patterns, whereas multicollinearity is a data structure issue that can actually increase model complexity without improving fit.

25
MCQhard

A healthcare company is developing a predictive model to identify patients at risk of readmission within 30 days. The data engineering team has built a pipeline that collects data from multiple sources, including electronic health records (EHR), lab results, and wearable device data. During initial testing, the model's performance is poor, with high false positives. Upon investigation, the team discovers that the data contains significant temporal misalignment: lab results are timestamped when ordered, not when collected; wearable data is aggregated hourly; and EHR data has inconsistent update frequencies. The data pipeline currently joins all features on the patient ID without aligning timestamps. The data volume is large, and processing time is a concern. Which action should the data engineering team take to most effectively address the issue and improve model performance?

A.Discard all records where timestamps do not match exactly across sources, and only use records with perfect alignment.
B.Implement a window-based feature aggregation (e.g., 6-hour windows) and align all features to the same time windows before joining.
C.Leave the pipeline unchanged and instead adjust the model's classification threshold to reduce false positives.
D.Use a data imputation algorithm to fill in missing timestamps and then join on the nearest timestamp.
AnswerB

Window-based aggregation aligns lab, wearable and EHR features onto shared 6-hour timestamps, removing the temporal misalignment that causes spurious correlations and false positives. It satisfies the processing-time constraint by aggregating incrementally rather than joining raw event-level records.

Why this answer

Temporal misalignment causes features to be joined at incorrect times, leading to data leakage or irrelevant features that degrade model performance. Implementing window-based aggregation aligns all features to consistent time windows (e.g., 6-hour) before joining, ensuring that features reflect the patient's state at the same point in time. This addresses the root cause and improves model accuracy.

Exam trap

AI0-001 often tests data preprocessing pitfalls; candidates might choose threshold adjustment or imputation as quick fixes, but the core issue is temporal alignment, which requires a systematic approach like windowing.

How to eliminate wrong answers

Option A is wrong because discarding records with misaligned timestamps would result in significant data loss and reduce the dataset size, potentially introducing bias and not solving the underlying issue of aligning features. Option C is wrong because adjusting the classification threshold only changes the decision boundary and does not fix the data quality problem; it may reduce false positives but at the cost of false negatives and does not improve the model's predictive power. Option D is wrong because imputing missing timestamps and joining on nearest timestamp can still introduce misalignment and may not accurately reflect the temporal relationships, especially with varying update frequencies.

26
MCQeasy

A data engineer needs to combine two datasets, each with unique customer_id, to include all records from both datasets. Which join type should be used?

A.FULL OUTER JOIN
B.RIGHT JOIN
C.LEFT JOIN
D.INNER JOIN
AnswerA

A FULL OUTER JOIN returns matched rows plus unmatched rows from both sides, padding missing columns with nulls. Since each dataset holds unique customer_id values, only this join preserves every record from both sources, as the stem requires.

Why this answer

A FULL OUTER JOIN returns all records from both datasets, matching rows where the customer_id is present in both and filling in NULLs for missing matches. This is the only join type that guarantees every unique customer_id from either dataset appears in the result, which is exactly what the requirement specifies.

Exam trap

CompTIA often tests the misconception that LEFT JOIN or RIGHT JOIN can include all records from both datasets, but candidates forget that these asymmetric joins exclude non-matching rows from the opposite side.

How to eliminate wrong answers

Option B (RIGHT JOIN) is wrong because it returns only all rows from the right dataset and matching rows from the left, omitting any customer_id that exists only in the left dataset. Option C (LEFT JOIN) is wrong because it returns only all rows from the left dataset and matching rows from the right, omitting any customer_id that exists only in the right dataset. Option D (INNER JOIN) is wrong because it returns only rows where customer_id exists in both datasets, discarding all non-matching records from either side.

27
Multi-Selecteasy

Which TWO data preprocessing techniques reduce the dimensionality of a dataset?

Select 2 answers
A.One-hot encoding
B.Imputation
C.Feature scaling
D.Principal Component Analysis (PCA)
E.Feature selection
AnswersD, E

PCA projects the original features onto orthogonal principal components ordered by explained variance, so the dataset can be represented with far fewer dimensions while retaining most information. This directly reduces dimensionality rather than merely selecting or transforming individual columns.

Why this answer

Principal Component Analysis (PCA) (D) is correct because it projects the original features onto a smaller set of orthogonal principal components that capture most of the variance, thereby reducing the number of dimensions while preserving as much information as possible. Feature selection (E) is also correct because it explicitly removes irrelevant or redundant features and keeps only a subset of the original variables, directly lowering the dataset's dimensionality. By contrast, one-hot encoding (A) increases dimensionality by creating a separate binary column per category, and imputation (B) merely fills in missing values without changing the number of features.

Feature scaling (C) standardizes or normalizes feature values (e.g., via min-max or z-score) but leaves the feature count unchanged, so it does not reduce dimensionality.

Exam trap

CompTIA often tests the distinction between techniques that transform or select features (reducing dimensionality) versus those that prepare data for modeling (like encoding, imputation, or scaling) without changing the number of features.

28
Multi-Selectmedium

Which THREE are common data preprocessing steps in a machine learning pipeline? (Choose 3)

Select 3 answers
A.Hyperparameter tuning
B.Encoding categorical variables
C.Model evaluation
D.Scaling numeric features
E.Handling missing values
AnswersB, D, E

Encoding categorical variables converts non-numeric labels into numeric representations, such as one-hot or ordinal encoding, so algorithms that require numerical input can process them. This satisfies the preprocessing requirement by making categorical features usable during model training, alongside handling missing values and feature scaling.

Why this answer

Encoding categorical variables (B) is a standard preprocessing step because algorithms require numeric input, so techniques like one-hot encoding or label encoding convert strings into numbers. Scaling numeric features (D) is also preprocessing, since methods such as standardization (z-score) or min-max normalization bring features to comparable ranges, which helps distance-based and gradient-based models. Handling missing values (E) is preprocessing too, using imputation (mean, median, mode) or deletion to ensure the dataset is complete before training.

Hyperparameter tuning (A) and model evaluation (C) are not preprocessing; they occur later in the pipeline, during model selection/training and after training, respectively.

Exam trap

CompTIA often tests the distinction between preprocessing steps (data cleaning, transformation) and later pipeline stages (model tuning, evaluation), so candidates mistakenly select hyperparameter tuning or model evaluation as preprocessing steps.

29
Multi-Selectmedium

A data scientist is preparing a dataset for training a machine learning model. The dataset contains a mix of numerical and categorical features, and some features have high cardinality. The data scientist needs to apply appropriate encoding techniques to transform categorical variables into a format suitable for the model. Which TWO encoding methods are most appropriate for high-cardinality categorical features? (Choose two.)

Select 2 answers
A.Target encoding
B.Label encoding
C.Frequency encoding
D.Binary encoding
E.One-hot encoding
AnswersA, C

Target encoding replaces each category with the mean of the target variable for that category. It is effective for high-cardinality features because it reduces dimensionality and captures the relationship between the category and the target. However, it can lead to overfitting if not properly regularized, especially with rare categories. Techniques like smoothing or adding noise can mitigate this. In this scenario, target encoding is a suitable method for high-cardinality categorical features.

Why this answer

Target encoding and frequency encoding are both effective for high-cardinality categorical features. Target encoding replaces categories with the mean target value, capturing predictive relationships, while frequency encoding replaces categories with their counts, reducing dimensionality without imposing order. One-hot encoding creates too many columns, label encoding imposes false order, and binary encoding may not capture target relationships as well.

Therefore, target and frequency encoding are the most appropriate methods.

Exam trap

The trap here is assuming that one-hot encoding is always the default for categorical variables, without considering the dimensionality explosion with high-cardinality features.

30
MCQmedium

A media company is building a recommendation model from user clickstream logs. The raw data arrives as millions of small JSON files in an object store, and nightly training jobs currently take over ten hours because the training cluster reads thousands of tiny files per second. The team wants to reduce training time without changing the model or the underlying data values. Which data engineering approach is most appropriate?

A.Convert the JSON files to CSV because CSV parsing is faster than JSON parsing.
B.Increase the number of worker nodes in the training cluster so more files can be read in parallel.
C.Cache the JSON files on local SSD storage on each worker node before training begins.
D.Compact the small JSON files into larger columnar files such as Parquet and have the training pipeline read those instead.
AnswerD

Consolidating many small JSON files into fewer larger Parquet files preserves all the underlying records while dramatically reducing per-file overhead, metadata operations, and list calls against the object store. Columnar layout also lets the training job read only the columns it needs and enables efficient compression, so the same data values feed the model with far less I/O time.

Why this answer

The bottleneck is the small-file pattern in object storage, not compute capacity or parsing speed. Repacking the same records into fewer large columnar files cuts listing, open, and read overhead and enables column pruning and compression, so the training job moves the same data values with far less I/O. This addresses the root cause while preserving the model and data semantics.

Exam trap

The trap here is assuming that adding more compute or switching file formats without consolidating files will fix an I/O-bound small-file problem.

31
Multi-Selectmedium

Which THREE are common causes of data leakage in machine learning pipelines?

Select 3 answers
A.Using time-based splitting for sequential data
B.Using future information to predict the present
C.Using cross-validation on the entire dataset
D.Applying normalization before splitting data into train and test sets
E.Including features that are directly derived from the target variable
AnswersB, D, E

Using future information to predict the present leaks target-correlated data backwards through time. In temporal pipelines, features computed from later events encode outcomes unavailable at prediction time, inflating validation scores while the deployed model cannot access that information.

Why this answer

Option B is correct because using future information to predict the present is the classic definition of temporal leakage: when training features contain values that would not have been available at prediction time (e.g., tomorrow's price used to predict today's), the model learns relationships that cannot generalize. Option D is correct because applying normalization (e.g., StandardScaler or MinMaxScaler) before splitting lets the scaler compute statistics such as the mean, variance, or min/max over the test data, so test-set information leaks into the training transformation; the scaler must be fit only on the training split. Option E is correct because features directly derived from the target variable (e.g., a 'remaining balance' column computed from the label, or target-encoded aggregates that include the current row's label) encode the answer into the inputs, producing artificially high validation scores that collapse in production.

Option A is not a leakage cause but a mitigation: time-based splitting is the recommended approach for sequential data precisely to prevent temporal leakage. Option C is also not inherently leakage: cross-validation on the entire dataset is standard practice as long as the preprocessing is fit within each fold; leakage arises only if transformations are fit on the full dataset before cross-validation.

Exam trap

CompTIA often tests the distinction between valid data splitting practices and actual leakage causes, so candidates may incorrectly select time-based splitting (Option A) as a leakage cause when it is actually a proper technique for sequential data.

32
MCQeasy

A data analyst is exploring a dataset and notices that one numerical feature has a highly skewed distribution with a long right tail. The analyst wants to apply a transformation to make the distribution more symmetric for a linear model. Which transformation is most appropriate?

A.Logarithmic transformation
B.Standardization (z-score normalization)
C.Square root transformation
D.Min-max normalization
AnswerA

A logarithmic transformation compresses the right tail and can make a right-skewed distribution more symmetric. It is effective when data spans several orders of magnitude and contains positive values. This helps linear models meet assumptions of normality and reduces the impact of outliers, improving model performance.

Why this answer

A logarithmic transformation is the most appropriate for a highly right-skewed distribution because it compresses large values and can make the distribution more symmetric. Square root is milder and less effective for high skewness, while standardization and min-max normalization only rescale without altering distribution shape.

Exam trap

The trap here is confusing scaling techniques like standardization or min-max normalization with transformations that actually change the distribution shape to reduce skewness.

33
MCQeasy

A data analyst is cleaning a dataset and finds that 20% of the values for the 'age' column are missing. Which imputation method is most robust if the data is not normally distributed?

A.Mean imputation
B.Median imputation
C.Mode imputation
D.Remove rows with missing values
AnswerB

The median resists skew and outliers because it depends on rank position rather than magnitude, so it stays representative when the distribution is non-normal. Mean imputation would distort the central tendency here, whereas the median satisfies the stem's non-normal constraint.

Why this answer

Median imputation is the most robust method for handling missing values in the 'age' column when the data is not normally distributed because the median is unaffected by outliers or skewness. Unlike the mean, which is sensitive to extreme values, the median provides a central tendency measure that better represents the typical value in non-normal distributions, preserving the dataset's integrity for downstream modeling.

Exam trap

CompTIA often tests the misconception that mean imputation is always the default or best choice for numerical data, but the trap here is that candidates overlook the importance of distribution shape and outlier sensitivity, leading them to select mean imputation despite the data not being normally distributed.

How to eliminate wrong answers

Option A is wrong because mean imputation assumes a normal distribution and is highly sensitive to outliers, which can introduce bias and distort the dataset's variance when the data is skewed. Option C is wrong because mode imputation is typically used for categorical data, not continuous variables like age, and it can lead to loss of granularity and inaccurate representation of the distribution. Option D is wrong because removing rows with missing values reduces sample size and can introduce selection bias, especially if the missingness is not completely at random, which is inefficient and may degrade model performance.

34
MCQmedium

A financial institution is training a risk assessment model. The dataset includes customer credit scores, income, age, and past loan defaults. During feature engineering, a data engineer creates a new feature 'income_to_debt_ratio'. Which type of feature engineering technique is this?

A.Feature encoding
B.Feature scaling
C.Feature selection
D.Feature combination
AnswerD

Feature combination creates new attributes by arithmetically relating two or more existing variables, here dividing income by debt to produce income_to_debt_ratio. This satisfies the stem's constraint that the engineer derived a single ratio from separate income and debt fields, rather than transforming one column or selecting a subset.

Why this answer

'income_to_debt_ratio' is created by combining two existing features (income and debt) into a single derived feature. This is a classic example of feature combination (also known as feature crossing or feature construction), where arithmetic operations or logical rules are applied to existing variables to generate new predictive signals. The goal is to capture interactions or relationships that the original features alone may not express linearly.

Exam trap

CompTIA often tests the distinction between feature engineering techniques by presenting a derived feature and expecting candidates to recognize it as feature combination rather than confusing it with scaling or encoding.

How to eliminate wrong answers

Option A is wrong because feature encoding transforms categorical variables into numerical representations (e.g., one-hot encoding, label encoding), not create new numerical ratios from existing numerical features. Option B is wrong because feature scaling normalizes or standardizes the range of feature values (e.g., min-max scaling, z-score normalization) without generating new features. Option C is wrong because feature selection reduces the number of features by choosing a subset of the original ones (e.g., using correlation analysis or recursive feature elimination), not by engineering new derived attributes.

35
MCQhard

A data scientist is building a model to detect anomalies in server logs. The dataset contains millions of log entries, each with a timestamp and a message. The scientist wants to create features that capture the frequency of certain keywords (e.g., 'error', 'timeout') over time. Which approach is MOST appropriate for creating these features while avoiding data leakage?

A.Compute keyword frequencies using a sliding window that only includes past log entries relative to each timestamp.
B.Compute keyword frequencies over the entire dataset and use them as static features.
C.Aggregate keyword frequencies per day and assign the same daily frequency to all entries within that day.
D.Use a bag-of-words representation for each log entry individually, ignoring timestamps.
AnswerA

Using a sliding window of past entries ensures that features are computed only from data available before the current timestamp, preventing leakage from future events. This mimics real-time detection where only historical data is available. It also captures temporal patterns like increasing error rates, which are useful for anomaly detection.

Why this answer

To avoid data leakage when creating time-based features, the computation must only use data available up to each point in time. A sliding window that aggregates past log entries ensures causality, mimicking real-time conditions. This captures temporal dynamics like increasing error rates, which are critical for anomaly detection, without using future information that would not be available in production.

Exam trap

The trap here is using global statistics or daily aggregates that inadvertently include future data, which artificially boosts model performance during training but fails in real-world deployment.

36
MCQmedium

A company is deploying an AI model to recommend products. The model's training data included historical purchases from the past two years, but the business environment has changed significantly due to a market shift. What is the most likely issue affecting model performance?

A.Concept drift
B.Overfitting
C.Underfitting
D.Data leakage
AnswerA

Concept drift occurs when the statistical relationship between input features and target changes over time. The market shift altered purchasing behaviour, so patterns learned from two-year-old data no longer reflect current reality, degrading recommendation relevance.

Why this answer

Concept drift occurs when the statistical properties of the target variable change over time, degrading model performance. In this scenario, the market shift alters customer purchasing patterns, making the historical training data (from the past two years) no longer representative of current behavior. This is the most likely issue because the model's recommendations will be based on outdated correlations.

Exam trap

The AI0-001 exam often tests the distinction between concept drift and data leakage, where candidates mistakenly attribute performance degradation to a data contamination issue rather than a shift in the underlying data distribution.

How to eliminate wrong answers

Option B is wrong because overfitting refers to a model that memorizes training data noise and fails to generalize, but the problem here is a change in the underlying data distribution, not excessive complexity. Option C is wrong because underfitting means the model is too simple to capture patterns in the training data, whereas the issue is that the training data itself no longer reflects the current environment. Option D is wrong because data leakage involves the accidental inclusion of future information in the training set, which is not described; the problem is a temporal shift in the data distribution, not a data contamination issue.

37
MCQmedium

A machine learning engineer is training a Support Vector Machine (SVM) with an RBF kernel on a dataset with features on different scales (e.g., age 0-100, income 0-1,000,000). The model converges slowly and yields poor accuracy. What should the engineer do first?

A.Standardize the features to have zero mean and unit variance
B.Increase the regularization parameter C to penalize misclassifications more
C.Decrease the gamma parameter to reduce the influence of each data point
D.Switch to a linear kernel to avoid distance calculations
AnswerA

The RBF kernel relies on Euclidean distance, so unscaled features such as income dominate age, distorting the kernel and slowing convergence. Standardising to zero mean and unit variance equalises feature influence, directly addressing the scale disparity.

Why this answer

Standardizing features to zero mean and unit variance is the correct first step because SVMs with RBF kernels are distance-based models. Features on vastly different scales (e.g., age 0-100 vs. income 0-1,000,000) cause the kernel to disproportionately weight larger-scale features, leading to slow convergence and poor accuracy. Standardization ensures each feature contributes equally to the distance calculations, improving both training speed and model performance.

Exam trap

The CompTIA AI+ exam often tests the misconception that hyperparameter tuning (C or gamma) is the primary fix for poor SVM performance, when in reality feature scaling is a prerequisite for distance-based kernels like RBF.

How to eliminate wrong answers

Option B is wrong because increasing the regularization parameter C penalizes misclassifications more heavily, which addresses overfitting or underfitting but does not fix the fundamental issue of feature scale disparity. Option C is wrong because decreasing the gamma parameter reduces the influence of each training point, which affects the decision boundary's smoothness but does not correct the scale imbalance that distorts distance computations. Option D is wrong because switching to a linear kernel avoids the RBF kernel's distance-based calculations but does not resolve the scale problem—linear SVMs also rely on distances and benefit from feature scaling.

38
MCQmedium

A dataset for a binary classification problem has 95% of samples in class "0" and 5% in class "1". The data scientist trains a logistic regression model and achieves 95% accuracy. Which metric should the scientist primarily use to evaluate model performance?

A.Precision, recall, and F1-score.
B.R-squared.
C.Accuracy.
D.Mean squared error.
AnswerA

With 95% class imbalance, a model predicting only class 0 scores 95% accuracy while detecting no positives. Precision, recall and F1-score expose this failure by measuring performance on the minority class rather than overall correctness.

Why this answer

In a highly imbalanced dataset (95% class 0, 5% class 1), accuracy is misleading because a model can achieve 95% accuracy by simply predicting the majority class for all samples. Precision, recall, and F1-score provide a more nuanced view of performance on the minority class, which is typically the class of interest in binary classification problems. The F1-score, in particular, balances precision and recall, making it the primary metric for evaluating model effectiveness on imbalanced data.

Exam trap

CompTIA often tests the concept that accuracy is a poor metric for imbalanced datasets, trapping candidates who assume high accuracy always indicates good model performance without considering class distribution.

How to eliminate wrong answers

Option B is wrong because R-squared is a metric for regression models, measuring the proportion of variance in the dependent variable explained by the independent variables, and is not applicable to classification tasks. Option C is wrong because accuracy is not a reliable metric for imbalanced datasets; a model that always predicts the majority class can achieve high accuracy without actually learning meaningful patterns, as seen with the 95% accuracy matching the class distribution. Option D is wrong because mean squared error (MSE) is a loss function for regression problems, used to quantify the average squared difference between predicted and actual continuous values, and is not appropriate for evaluating binary classification outputs.

39
MCQeasy

A team is building a regression model to predict house prices. Which data transformation is most appropriate if the target variable exhibits right skewness?

A.Principal component analysis (PCA)
B.Standardization (Z-score)
C.One-hot encoding
D.Log transformation
AnswerD

Log transformation compresses the long right tail of a positively skewed target, pulling extreme values toward the mean and stabilising variance. This satisfies the right-skewness constraint, making the target closer to normal so linear regression's residual assumptions hold and predictions are less distorted.

Why this answer

Log transformation is the most appropriate technique for right-skewed target variables because it compresses the long tail, making the distribution more symmetric and closer to Gaussian. This stabilizes variance and often improves the performance of regression models that assume normally distributed errors, such as linear regression.

Exam trap

CompTIA often tests the misconception that standardization can fix skewness, but candidates must remember that standardization only rescales the data, not reshape its distribution.

How to eliminate wrong answers

Option A is wrong because Principal Component Analysis (PCA) is a dimensionality reduction technique for features, not a transformation applied to the target variable; it does not address skewness in the target. Option B is wrong because Standardization (Z-score) centers and scales the data but does not change the shape of the distribution, so it cannot correct right skewness. Option C is wrong because One-hot encoding is used to convert categorical variables into numerical format, not to transform a continuous target variable.

40
MCQmedium

A data engineer is preparing a dataset for a machine learning model that predicts customer lifetime value. The dataset contains a 'monthly_spend' column with a highly right-skewed distribution. The team wants to apply a transformation to reduce skewness while preserving the relative order of values. Which transformation is most appropriate?

A.Standardize the 'monthly_spend' column by subtracting the mean and dividing by the standard deviation.
B.Replace each value with its rank among all values in the column (rank transformation).
C.Apply one-hot encoding to the 'monthly_spend' column by binning into deciles.
D.Apply a logarithmic transformation (log(1 + x)) to the 'monthly_spend' column.
AnswerD

A logarithmic transformation compresses the scale of large values while maintaining the order of observations, effectively reducing right skewness. Using log(1+x) handles zero values safely. This is a standard technique for monetary data with positive skew, and it helps linear models and neural networks learn more effectively by making the distribution closer to normal without losing rank information.

Why this answer

A logarithmic transformation is the most appropriate because it compresses large values and reduces right skewness while preserving the relative order of observations. Unlike standardization, it changes the distribution shape; unlike binning or rank transformation, it retains the continuous nature and magnitude relationships, making it well-suited for monetary data that is positively skewed.

Exam trap

The trap here is assuming that standardization or normalization alone can correct skewness, when in fact they only rescale without altering the distribution's shape.

41
MCQmedium

A data scientist is building a model to predict the likelihood of a customer defaulting on a loan. The dataset contains a feature 'debt_to_income_ratio' that is highly skewed with a long right tail. The scientist decides to apply a logarithmic transformation to this feature. Which statement best describes the effect of this transformation?

A.It normalizes the feature to a range between 0 and 1, which is required for all machine learning algorithms.
B.It converts the feature into a categorical variable, which is useful for tree-based models.
C.It reduces the impact of outliers and makes the distribution more symmetric, which can help linear models perform better.
D.It increases the variance of the feature, allowing the model to capture more complex patterns.
AnswerC

A logarithmic transformation compresses the scale of large values, pulling in the right tail and making the distribution more normal-like. This reduces the influence of extreme outliers and can improve the performance of linear models that assume normally distributed errors or linear relationships.

Why this answer

Log transformation is commonly used to handle right-skewed data. It reduces skewness, stabilizes variance, and makes the distribution more symmetric, which can improve the performance of linear models and neural networks that are sensitive to feature scaling. It does not increase variance, convert to categorical, or normalize to a fixed range.

Exam trap

The trap here is confusing log transformation with normalization or standardization, leading to the incorrect belief that it scales features to a specific range.

42
MCQhard

An AI model is deployed to a mobile app with limited computational resources. The model is a deep neural network with high latency. Which technique is best to reduce inference time?

A.Increase batch size
B.Add more layers
C.Use a larger model
D.Quantization
AnswerD

Quantization converts weights and activations from 32-bit floats to 8-bit integers, shrinking memory footprint and enabling faster integer arithmetic on resource-constrained mobile hardware. This directly cuts inference latency while preserving acceptable accuracy, satisfying the limited-compute constraint.

Why this answer

Quantization reduces the precision of the model's weights and activations (e.g., from 32-bit floating point to 8-bit integer), which decreases memory footprint and speeds up computation on resource-constrained devices like mobile phones. This directly lowers inference latency without requiring additional hardware or architectural changes.

Exam trap

The AI0-001 exam often tests the misconception that increasing batch size or model size improves performance on edge devices, when in fact these techniques increase resource demands and latency in low-resource environments.

How to eliminate wrong answers

Option A is wrong because increasing batch size improves throughput (samples per second) but does not reduce per-sample latency; it actually increases memory usage and can worsen latency on mobile devices with limited resources. Option B is wrong because adding more layers increases the model depth, which increases computational complexity and latency, making inference slower. Option C is wrong because using a larger model (more parameters) increases both memory and compute requirements, directly increasing inference time on constrained devices.

43
MCQmedium

A data engineer is designing a data pipeline for a machine learning model that predicts equipment failure in a manufacturing plant. The pipeline ingests sensor data every second, and the model must be retrained daily with the latest data. The team needs to ensure that the training data is representative of the current operating conditions and that the model does not become stale. Which strategy is most appropriate for maintaining the training dataset?

A.Implement a sliding window approach where the training dataset includes only the most recent 30 days of sensor data.
B.Randomly sample 10% of all historical data for each daily retraining to reduce computational load.
C.Use the entire historical dataset from the plant's inception to train the model, ensuring maximum data volume.
D.Manually curate a fixed dataset of known failure events and use it for all retrainings.
AnswerA

A sliding window ensures that the training data reflects the most recent operating conditions, which is crucial when equipment behavior or environmental factors change over time. By using only the last 30 days, the model adapts to concept drift and remains relevant. This approach balances the need for sufficient data with the need for recency, and it is a common practice in time-series forecasting and predictive maintenance.

Why this answer

A sliding window of recent data is most appropriate because it keeps the training set aligned with current operating conditions, mitigating concept drift. This ensures the model learns from the most relevant examples, which is essential for predictive maintenance where equipment behavior evolves over time.

Exam trap

The trap here is assuming that more data is always better, but in dynamic environments, outdated data can actually harm model performance by introducing irrelevant patterns.

44
MCQeasy

A model's training accuracy is 99% but validation accuracy drops to 60%. What is the most likely issue?

A.Data leakage
B.Overfitting
C.Multicollinearity
D.Underfitting
AnswerB

A large gap between 99% training accuracy and 60% validation accuracy means the model memorised training noise rather than generalising. Overfitting is the mechanism: high variance causes the model to fit patterns specific to the training set that do not hold on unseen validation data.

Why this answer

A training accuracy of 99% with a validation accuracy of only 60% is a classic symptom of overfitting. The model has memorized the training data, including noise and outliers, rather than learning generalizable patterns, causing it to perform poorly on unseen validation data.

Exam trap

CompTIA often tests the distinction between overfitting and data leakage by presenting a large accuracy gap, where candidates might mistakenly attribute the issue to data leakage instead of recognizing that leakage typically inflates both accuracies rather than creating a divergence.

How to eliminate wrong answers

Option A is wrong because data leakage typically causes both training and validation accuracy to be artificially high, not a large gap between them; it occurs when information from outside the training set inadvertently influences the model. Option C is wrong because multicollinearity refers to high correlation among input features in regression models, which affects coefficient stability and interpretability, not a drastic accuracy drop between training and validation sets. Option D is wrong because underfitting would result in low accuracy on both training and validation sets (e.g., both below 70%), not a high training accuracy with a low validation accuracy.

45
MCQeasy

A data engineer needs to store training data in a format that supports columnar pruning during model training. Which storage format should they use?

A.Parquet
B.XML
D.CSV
AnswerA

Parquet stores data column-wise with per-column statistics and row-group metadata, letting readers skip irrelevant columns and row groups entirely. This columnar pruning reduces I/O during training, directly meeting the stem's requirement for efficient columnar access.

Why this answer

Parquet is the correct choice because it is a columnar storage format that enables column pruning, allowing the training process to read only the columns needed for model training rather than entire rows. This reduces I/O and speeds up data loading, which is critical for large-scale AI/ML workloads. Unlike row-oriented formats, Parquet stores data by columns, making it efficient for analytical queries and feature selection.

Exam trap

CompTIA often tests the misconception that JSON or CSV are acceptable for columnar pruning because they are common and human-readable, but the trap here is that only columnar formats like Parquet or ORC support efficient column-level access, while row-oriented formats require full record scans.

How to eliminate wrong answers

Option B (XML) is wrong because XML is a verbose, hierarchical text format that stores data row-wise and lacks columnar pruning capabilities, leading to high storage overhead and slow read performance for tabular data. Option C (JSON) is wrong because JSON is a row-oriented, self-describing format that requires parsing entire records even when only a subset of fields is needed, making it unsuitable for column pruning. Option D (CSV) is wrong because CSV is a flat, row-oriented text format that forces reading entire rows into memory, with no support for columnar storage or predicate pushdown, resulting in inefficient I/O for selective column access.

46
MCQhard

A fraud detection team trains a gradient boosted tree model on transaction data. During evaluation, the team notices the model performs extremely well on the training set but poorly on a holdout set drawn from the same time period. Investigation shows that a feature named 'chargeback_flag' is populated only after a dispute is resolved, sometimes weeks after the transaction. The team wants to deploy the model to score transactions in real time. Which action best addresses the problem?

A.Increase the size of the holdout set so the evaluation becomes more statistically reliable.
B.Replace the flag with a rolling average of the customer's past chargebacks to preserve some of its signal.
C.Remove the 'chargeback_flag' feature and retrain the model using only features that are available at transaction scoring time.
D.Apply stronger regularization and reduce the model's maximum depth to prevent it from relying on the flag.
AnswerC

The flag is a label leak: it is recorded only after the outcome the model is supposed to predict is known, so it cannot exist when scoring a new transaction. Removing it and retraining on features that are genuinely available at inference time eliminates the leak and produces a model whose offline metrics better reflect real-time performance. This is the correct root-cause fix.

Why this answer

The 'chargeback_flag' is populated only after a dispute is resolved, which is after the transaction outcome is known, so it leaks the label into training. The correct fix is to identify features that are actually available at real-time scoring and retrain without the leaky field. Hyperparameter tuning, larger validation sets, or partial replacements do not remove the leak and will keep offline metrics misleadingly high.

Exam trap

The trap here is treating high offline accuracy as a modeling problem to tune rather than recognizing that a post-outcome feature is leaking the label.

47
Multi-Selectmedium

A data scientist is building a model to predict equipment failure using sensor data. The dataset contains time-series readings from multiple sensors, and the goal is to detect anomalies that precede failures. Which TWO feature engineering techniques are most appropriate for this time-series data? (Choose two.)

Select 2 answers
A.Replace missing sensor values with the overall mean of the entire dataset.
B.Apply one-hot encoding to the timestamp column to represent each time point as a binary vector.
C.Compute rolling window statistics such as mean, standard deviation, and min/max over recent time intervals.
D.Perform principal component analysis (PCA) on the raw sensor readings to reduce dimensionality.
E.Extract lag features by including previous sensor readings as additional input variables.
AnswersC, E

Rolling window statistics capture temporal patterns and trends, such as increasing variance before failure. They summarize recent behavior and are effective features for anomaly detection in sensor data. These features help models identify deviations from normal operating conditions, improving predictive performance.

Why this answer

Rolling window statistics and lag features are essential for time-series data because they encode temporal dependencies and trends. Rolling statistics summarize recent behavior, while lag features provide historical context. Together, they enable the model to detect anomalies that precede equipment failure.

The other options either ignore time order or are not suitable for temporal feature extraction.

Exam trap

The trap here is selecting generic dimensionality reduction or imputation methods that ignore the sequential nature of time-series data, rather than techniques that explicitly capture temporal patterns.

48
MCQhard

A large e-commerce company uses a recommendation system based on collaborative filtering. The system uses a matrix factorization model that is trained nightly on the entire user-item interaction history. Recently, the company launched a flash sale with thousands of new products. Users are reporting that the recommendations are not showing the new products, even for users who have purchased them during the sale. The data engineering team notices that the new products have very few interactions in the training data. The model's loss on the validation set has increased, and the recall@10 metric has dropped from 0.45 to 0.32. The team needs to improve the recommendation of new items without retraining the entire model from scratch every hour. Which approach should the team take?

A.Use a hybrid model that combines collaborative filtering with content-based features from product metadata
B.Retrain the model every hour to incorporate new interactions quickly
C.Remove the new products from the recommendation pool until they accumulate enough interactions
D.Increase the number of latent factors in the matrix factorization model
AnswerA

Content-based metadata features let the hybrid model score new products via their attributes rather than interaction counts, solving the cold-start recall drop without hourly full retraining. Matrix factorisation alone cannot represent items with few interactions.

Why this answer

A hybrid model that combines collaborative filtering with content-based features (e.g., product metadata like category, price, or description) can recommend new products even with zero or very few user interactions. The content-based component leverages item attributes to compute similarity between new and existing items, enabling the system to surface new products without requiring extensive interaction history. This approach addresses the cold-start problem for new items while preserving the collaborative filtering signal for established items, and it does not require retraining the entire model from scratch every hour.

Exam trap

CompTIA often tests the misconception that simply retraining more frequently or increasing model complexity (e.g., more latent factors) can solve the cold-start problem, but the core issue is the lack of interaction data for new items, which requires a content-based or hybrid approach to leverage item metadata.

How to eliminate wrong answers

Option B is wrong because retraining the model every hour would be computationally expensive and operationally impractical for a large-scale system with thousands of new products; it also does not solve the fundamental cold-start issue since new items still have very few interactions in each hourly training window. Option C is wrong because removing new products from the recommendation pool defeats the business purpose of the flash sale, which is to promote and surface new items to users, and it would lead to a poor user experience and lost revenue. Option D is wrong because increasing the number of latent factors in matrix factorization does not address the lack of interaction data for new items; it may even exacerbate overfitting to sparse data and increase computational cost without improving cold-start recommendations.

49
Multi-Selecthard

Which TWO are best practices for versioning machine learning models? (Choose 2)

Select 2 answers
A.Use the same model version for all deployments
B.Tag each model with training date, hyperparameters, and performance metrics
C.Use a version control system (e.g., Git) for model code and configuration
D.Store only the final model binary without metadata
E.Manually rename model files with version numbers
AnswersB, C

Recording training date, hyperparameters, and performance metrics alongside each model creates a reproducible audit trail, letting teams trace which configuration produced which result and compare candidates. This metadata satisfies the traceability requirement that versioning practices demand.

Why this answer

Option B is correct because tagging each model with its training date, hyperparameters, and performance metrics creates an auditable lineage that lets teams reproduce results, compare candidates, and roll back to a known-good model when production metrics degrade. Option C is correct because placing model code and configuration under a version control system such as Git provides immutable commit history, branching, code review, and the ability to correlate a deployed artifact with the exact source revision that produced it. Together, B and C satisfy the core ML versioning requirements of reproducibility, traceability, and governance.

Option A is wrong because reusing one model version across all deployments eliminates the ability to distinguish, roll back, or A/B test different models. Option D is wrong because storing only the final binary without metadata makes the model impossible to reproduce, audit, or troubleshoot. Option E is wrong because manually renaming files is error-prone, unauditable, and does not capture training context or enable automated deployment pipelines.

Exam trap

CompTIA often tests the misconception that versioning is only about file naming or storing the binary, when in fact it requires a comprehensive metadata and code tracking system to ensure reproducibility and traceability.

50
MCQmedium

A machine learning engineer is preparing a dataset for a model that predicts whether a customer will click on an ad. The dataset contains a feature 'time_since_last_purchase' measured in hours, which has a highly skewed distribution with a long tail. The engineer decides to apply a logarithmic transformation to this feature. Which statement BEST describes the effect of this transformation?

A.It converts the feature into a categorical variable by binning values into logarithmic intervals.
B.It normalizes the feature to a 0-1 range, ensuring all features contribute equally to distance calculations.
C.It reduces the impact of extreme values and makes the distribution more symmetric, which can help linear models.
D.It eliminates the need for feature scaling because the transformed values are already standardized.
AnswerC

A logarithmic transformation compresses the range of large values, reducing skewness and making the distribution closer to normal. This can improve the performance of linear models that assume normally distributed features and are sensitive to outliers. It also stabilizes variance, which is beneficial for models like linear regression or logistic regression.

Why this answer

Applying a logarithmic transformation to a skewed feature like time since last purchase compresses the long tail and reduces skewness. This makes the feature more symmetric, which can improve the performance of models that assume normality or are sensitive to outliers, such as linear models. It does not normalize to a fixed range, bin the feature, or standardize it; those are separate preprocessing steps.

Exam trap

The trap here is confusing log transformation with normalization or standardization, and assuming it alone makes features comparable without additional scaling.

51
MCQeasy

An e-commerce company deploys a model to recommend products to users. The recommendation system uses collaborative filtering based on user-item interaction history. After deployment, the model shows decreasing click-through rates (CTR) over time. The data engineer notices that the model was trained on data from the past six months and is retrained daily. However, the trend suggests that user preferences are shifting more rapidly than expected. The engineer suspects that the model is suffering from distribution drift. Which approach should the engineer implement to adapt the model more quickly to changing user behavior?

A.Increase the retraining period to once per week to reduce computational cost
B.Switch to an online learning algorithm that updates the model after each user click
C.Increase the model complexity by adding more features and layers
D.Use only the last week of data for training to focus on recent trends
AnswerB

Online learning updates weights incrementally after each click, so the model tracks rapidly shifting user preferences between daily retrains. This directly addresses the distribution drift constraint, where batch retraining on six-month historical data lags behind the observed CTR decline.

Why this answer

Switching to an online learning algorithm that updates the model after each user click allows the recommendation system to adapt in near real-time to shifting user preferences, directly addressing the rapid distribution drift. This is the most responsive approach when preferences change faster than daily retraining can capture.

Exam trap

The trap is choosing a batch-oriented fix (more data, more features, different retraining cadence) when the problem is the speed of adaptation — candidates must recognize that only an online or incremental learning approach can keep pace with rapid distribution drift.

How to eliminate wrong answers

Option A is wrong because increasing the retraining period to weekly would make the model even slower to adapt, worsening the drift problem. Option C is wrong because increasing model complexity does not address distribution drift — a more complex model trained on stale data will still be stale, and may overfit. Option D is wrong because using only the last week of data may help slightly but still relies on daily batch retraining, which is too slow for rapidly shifting preferences and may discard valuable longer-term signals.

52
Multi-Selectmedium

Which TWO techniques are commonly used for feature selection in machine learning? (Choose 2)

Select 2 answers
A.Principal Component Analysis (PCA)
B.SMOTE
C.L1 regularization (Lasso)
D.Dropout
E.Recursive Feature Elimination (RFE)
AnswersC, E

L1 regularization adds a penalty proportional to the absolute value of coefficients, driving irrelevant feature weights exactly to zero and thereby performing embedded feature selection. This yields a sparse model, satisfying the feature-selection technique requirement rather than merely shrinking coefficients as L2 does.

Why this answer

L1 regularization (Lasso) is correct because it adds a penalty equal to the absolute value of the coefficients to the loss function, which drives less important feature coefficients exactly to zero, effectively performing embedded feature selection. Recursive Feature Elimination (RFE) is correct because it is a wrapper method that repeatedly trains a model, ranks features by importance (e.g., coefficients or feature_importances_), removes the weakest feature(s), and recurses until the desired number of features remains. PCA is not a feature-selection technique but a dimensionality-reduction method that creates new uncorrelated components from the original features, so it does not select among them.

SMOTE is a data-level technique for handling class imbalance by synthesizing minority-class samples, not for selecting features. Dropout is a neural-network regularization method that randomly deactivates neurons during training to reduce overfitting, and it does not perform feature selection.

Exam trap

CompTIA often tests the distinction between dimensionality reduction (PCA) and feature selection, where candidates mistakenly think PCA selects original features rather than creating new ones.

53
MCQhard

A deep learning model for image classification achieves 99% training accuracy but only 85% validation accuracy. The model has millions of parameters. Which technique is most likely to reduce overfitting while maintaining high accuracy?

A.Reduce batch size from 32 to 8
B.Decrease the learning rate by a factor of 10
C.Add dropout layers with a rate of 0.5 after each convolutional block
D.Increase the number of training epochs to 500
AnswerC

Dropout randomly deactivates neurons during training, preventing co-adaptation and forcing distributed representations. A rate of 0.5 after each convolutional block substantially regularises the millions of parameters, narrowing the training-validation gap while retaining capacity for high accuracy.

Why this answer

Dropout is a regularization technique that randomly drops a fraction of neurons during training, which prevents co-adaptation of features and forces the network to learn more robust representations. With 99% training accuracy and 85% validation accuracy, the model is clearly overfitting, and adding dropout layers with a rate of 0.5 after each convolutional block directly addresses this by reducing the model's capacity to memorize the training data, while still allowing high accuracy on the validation set.

Exam trap

The AI0-001 exam often tests the misconception that reducing learning rate or batch size is a primary method to combat overfitting, when in fact these are optimization adjustments, not regularization techniques designed to reduce model capacity.

How to eliminate wrong answers

Option A is wrong because reducing batch size from 32 to 8 increases gradient noise and can actually lead to slower convergence or instability, but it does not directly regularize the model to combat overfitting; it may even worsen generalization in some cases. Option B is wrong because decreasing the learning rate by a factor of 10 helps with convergence and fine-tuning but does not address the root cause of overfitting—it only changes the step size, not the model's capacity to memorize. Option D is wrong because increasing the number of training epochs to 500 will only exacerbate overfitting, as the model will have more iterations to fit the training data noise, likely driving validation accuracy even lower.

54
MCQeasy

An engineer is building a regression model to predict housing prices. The dataset includes features such as square footage, number of bedrooms, and year built. The engineer notices that the square footage values range from 500 to 10,000, while the number of bedrooms ranges from 1 to 5. Which preprocessing step is most critical before training a gradient descent-based model?

A.Use k-fold cross-validation
B.Apply log transformation to all features
C.Normalize or standardize the features
D.One-hot encode the features
AnswerC

Square footage spans 500–10,000 while bedrooms span 1–5; this scale disparity makes gradient descent oscillate, as the larger-range feature dominates the loss surface. Standardising or normalising features places them on comparable scales, accelerating convergence and satisfying the stem's preprocessing requirement for a gradient descent-based model.

Why this answer

Gradient descent-based models are sensitive to the scale of input features because they update weights proportionally to the gradient, which is influenced by feature magnitudes. With square footage ranging 500–10,000 and bedrooms 1–5, the larger feature will dominate the gradient, causing slow or unstable convergence. Normalizing or standardizing (e.g., Z-score or min-max scaling) ensures all features contribute equally, leading to faster and more reliable training.

Exam trap

CompTIA often tests the misconception that any data transformation (like log or one-hot encoding) is universally beneficial, but the key is matching the preprocessing step to the model's mathematical requirements—here, gradient descent's sensitivity to scale makes normalization/standardization the critical step.

How to eliminate wrong answers

Option A is wrong because k-fold cross-validation is a model evaluation technique to assess generalization, not a preprocessing step to address feature scale issues. Option B is wrong because log transformation is used to handle skewed distributions or multiplicative relationships, not to rescale features with different ranges; applying it to all features (including integer counts like bedrooms) can distort their meaning and is unnecessary for gradient descent scaling. Option D is wrong because one-hot encoding is used for categorical features to convert them into binary vectors, but the features listed (square footage, bedrooms, year built) are all numerical and do not require encoding.

55
Multi-Selecthard

A data engineer is designing a pipeline for a streaming data application that uses a machine learning model to detect anomalies in real time. Which TWO practices should the engineer implement to ensure data quality and model reliability?

Select 2 answers
A.Use batch processing to transform data in fixed intervals
B.Store all raw data indefinitely for future analysis
C.Use a sliding window for feature computation
D.Implement data validation checks at the ingestion point
E.Retrain the model on a fixed schedule every 24 hours
AnswersC, D

A sliding window recomputes features over the most recent events, keeping anomaly scores aligned with current behaviour. Fixed or expanding windows let stale data dilute the signal, degrading real-time detection accuracy as stream characteristics drift.

Why this answer

Option C is correct because a sliding window computes features over a continuously advancing time range, which is essential for real-time anomaly detection since it captures recent, temporally relevant data points while maintaining feature freshness as the stream evolves. Option D is correct because implementing data validation checks at the ingestion point catches malformed, missing, or out-of-range records before they reach the model, preventing garbage-in-garbage-out failures and preserving both data quality and model reliability in a streaming pipeline. Option A is not appropriate because batch processing in fixed intervals introduces latency and defeats the real-time requirement of the streaming application.

Option B is not required for data quality or model reliability; storing all raw data indefinitely is a retention/archival decision and does not itself validate or improve streaming data. Option E is not ideal because a fixed 24-hour retraining schedule cannot adapt to concept drift or anomalies that emerge within the stream, and retraining cadence should be driven by drift detection rather than an arbitrary fixed interval.

Exam trap

CompTIA often tests the misconception that batch processing or fixed retraining schedules are sufficient for real-time streaming applications, when in fact sliding windows and continuous validation are required to maintain low latency and model accuracy.

56
MCQhard

A data pipeline ingests streaming data from IoT sensors. The current batch processing pipeline causes stale predictions. Which architecture change is most appropriate?

A.Use a larger batch interval
B.Revert to micro-batch processing with Apache Spark
C.Store raw data in Hadoop HDFS
D.Implement Apache Kafka and stream processing
AnswerD

Apache Kafka decouples ingestion from processing and supports continuous stream processing, so events are handled as they arrive rather than accumulated into batches. This eliminates the latency that caused stale predictions, satisfying the requirement for near-real-time output from IoT sensor data.

Why this answer

Apache Kafka plus a stream processing engine (e.g., Kafka Streams, Flink, or Spark Structured Streaming) processes events as they arrive, eliminating the staleness inherent in batch pipelines. This directly addresses the requirement to move from stale batch predictions to real-time ingestion and inference. It is the only option that changes the architecture from batch to true streaming.

Exam trap

AI0-001 often tests the misconception that micro-batch is 'real-time' — candidates pick Spark micro-batch because it sounds streaming, missing that true stream processing (Kafka/Flink) is required for low-latency predictions.

How to eliminate wrong answers

Option A is wrong because increasing the batch interval makes predictions even staler, worsening the problem. Option B is wrong because micro-batch with Spark is still batch-oriented and introduces latency proportional to the micro-batch interval; it does not provide continuous stream processing. Option C is wrong because storing raw data in HDFS is a storage decision that does nothing to reduce prediction latency — HDFS is batch-oriented and adds no streaming capability.

57
Multi-Selectmedium

A data engineer is designing a feature store for machine learning. Which THREE components are essential for a feature store? (Choose THREE.)

Select 3 answers
A.Data ingestion pipeline
B.Online serving layer
C.Feature repository
D.Experiment tracking
E.Model registry
AnswersA, B, C

Ingestion pipelines move raw data from source systems into the feature store, populating both offline and online stores. Without this component, features cannot be created, refreshed or kept consistent, so it is fundamental to any feature store architecture.

Why this answer

A feature store fundamentally requires a data ingestion pipeline (A) to pull raw data from sources such as batch files, streaming events, or databases and transform it into features, since without ingestion there would be no features to store or serve. The online serving layer (B) is essential because it provides low-latency access to the latest feature values for real-time inference, typically backed by a low-latency store like Redis or DynamoDB. The feature repository (C) is also essential as the central catalog that stores feature definitions, metadata, and versioned feature values, enabling reuse and consistency between training and serving.

Experiment tracking (D) and model registry (E) are important MLOps components but belong to the model development and deployment lifecycle, not to the core architecture of a feature store, so they are not essential components of a feature store itself.

Exam trap

CompTIA AI often tests candidates by including components from the broader ML lifecycle (like experiment tracking and model registry) to distract from the specific, essential components of a feature store.

58
MCQmedium

A machine learning model for credit card fraud detection is deployed. The model's precision is 0.95 and recall is 0.60. The business cost of missing a fraud is very high. Which of the following should the team prioritize to reduce the number of false negatives?

A.Use a different model algorithm.
B.Add more features.
C.Increase the classification threshold.
D.Decrease the classification threshold.
AnswerD

Lowering the classification threshold makes the model classify more cases as fraudulent, directly increasing recall and reducing false negatives — the costly misses the stem prioritises. Precision will fall as more legitimate transactions are flagged, but that trade-off is acceptable when the business cost of undetected fraud is very high.

Why this answer

Decreasing the classification threshold makes the model more sensitive, classifying more transactions as fraudulent. This increases recall (reducing false negatives) at the cost of precision. Given the high cost of missing fraud, lowering the threshold is the direct way to capture more true positives, even if it increases false positives.

Exam trap

CompTIA often tests the misconception that improving model accuracy or changing algorithms is the primary fix, when in fact adjusting the decision threshold is the simplest and most effective way to address precision-recall trade-offs for high-cost false negatives.

How to eliminate wrong answers

Option A is wrong because simply switching algorithms does not guarantee a reduction in false negatives; the threshold and cost function matter more. Option B is wrong because adding more features may improve overall model performance but does not directly target the trade-off between precision and recall; it could even increase false negatives if the new features are noisy. Option C is wrong because increasing the classification threshold makes the model more conservative, reducing false positives but increasing false negatives, which is the opposite of what is needed.

59
Multi-Selectmedium

A data engineer is building a pipeline to ingest and process data from various sources for an AI model. The pipeline must handle both structured data from relational databases and unstructured text from documents. The engineer needs to ensure data quality and prepare the data for model training. Which TWO actions are MOST appropriate for handling missing values in the structured data? (Choose two.)

Select 2 answers
A.Impute missing numerical values with the median of the available data.
B.Remove all rows with any missing values to ensure a complete dataset.
C.Encode missing values as a separate category for numerical features.
D.Use a constant value like zero to fill missing numerical values.
E.Apply a model-based imputation method such as k-nearest neighbors (KNN) imputation.
AnswersA, E

Imputing with the median is a robust method for numerical features because it is less sensitive to outliers than the mean. It preserves the central tendency and allows the model to use all samples. This is a common and effective approach when missingness is random and the proportion of missing data is not too high.

Why this answer

For structured data with missing numerical values, median imputation is a robust simple method, while KNN imputation leverages feature correlations for more accurate estimates. Both preserve data and avoid the biases of deletion or arbitrary constant filling. These actions are appropriate for preparing data for model training, ensuring quality and completeness without introducing severe distortions.

Exam trap

The trap here is assuming that removing all rows with missing values is always safe or that filling with zero is harmless, when both can introduce bias and degrade model performance.

60
MCQhard

Refer to the exhibit. A data scientist reviews the MLflow run for a Random Forest model on customer churn data. What is the most likely issue with this model?

A.The model is underfitting because training accuracy is too high.
B.The model is overfitting because there is a large gap between train and validation accuracy.
C.The model is performing well because validation accuracy is above 0.8.
D.The model has a data leak because dataset version is v2.
AnswerB

A wide gap between training and validation accuracy is the diagnostic signature of high variance: the Random Forest has memorised training noise rather than generalising. That gap, not the absolute accuracy, is what identifies overfitting in the MLflow run for this churn model.

Why this answer

A large gap between training accuracy (e.g., 0.99) and validation accuracy (e.g., 0.82) indicates that the Random Forest model has memorized the training data but fails to generalize to unseen validation data. This is the classic symptom of overfitting, where the model captures noise rather than the underlying pattern. In MLflow, comparing train and validation metrics directly reveals this discrepancy.

Exam trap

CompTIA often tests the misconception that high validation accuracy alone indicates a good model, ignoring the critical comparison between training and validation metrics to detect overfitting.

How to eliminate wrong answers

Option A is wrong because underfitting is characterized by low training accuracy, not high training accuracy; high training accuracy with poor validation performance indicates overfitting, not underfitting. Option C is wrong because a validation accuracy above 0.8 alone does not guarantee good model performance if there is a significant gap between train and validation accuracy, which signals overfitting. Option D is wrong because dataset version v2 is simply a versioning label and does not inherently cause data leakage; data leakage would involve information from the validation set leaking into training, which is unrelated to the version number.

61
MCQmedium

During model deployment, a data engineer notices that the model's predictions are consistently lower than expected due to a shift in the distribution of one feature between training and production. Which technique should be used to detect and quantify this shift?

A.Compute the root mean square error (RMSE)
B.Calculate the population stability index (PSI)
C.Generate a confusion matrix
D.Perform a t-test on the means
AnswerB

Calculating the population stability index directly quantifies distributional drift between training and production data for a single feature, satisfying the stem's requirement to detect and measure the shift. PSI compares binned proportions across both datasets, producing a numeric score that flags meaningful divergence, unlike accuracy metrics which cannot isolate feature-level change.

Why this answer

The Population Stability Index (PSI) is specifically designed to detect and quantify shifts in the distribution of a feature or score between two populations, such as training and production datasets. It measures the stability of the feature by comparing the proportion of observations in each bin across the two time periods, making it the correct choice for diagnosing distribution drift in model deployment.

Exam trap

CompTIA often tests the distinction between performance metrics (like RMSE or confusion matrix) and distribution monitoring metrics (like PSI), trapping candidates who confuse model accuracy evaluation with data drift detection.

How to eliminate wrong answers

Option A is wrong because RMSE measures the average magnitude of prediction errors, not distribution shifts between datasets. Option C is wrong because a confusion matrix evaluates classification performance against ground truth labels, not feature distribution changes. Option D is wrong because a t-test on the means only checks for a difference in central tendency, not the full distributional shift that PSI captures, and it is sensitive to sample size rather than bin-wise stability.

62
MCQmedium

A retail company is building a recommendation system to suggest products to customers based on their purchase history. The data engineering team has collected data from point-of-sale systems, online browsing logs, and customer reviews. After cleaning the data, they notice that the feature set has over 500 dimensions, leading to high computational costs and potential overfitting. They need to reduce dimensionality while preserving as much variance as possible for the model. The team is considering various techniques. Which approach should they take to achieve this goal most effectively?

A.Keep all features but apply L1 regularization (Lasso) in the model to automatically reduce coefficients to zero.
B.Apply t-Distributed Stochastic Neighbor Embedding (t-SNE) to reduce the feature space to 50 dimensions.
C.Select only features that have a high correlation with the target variable, discarding all others.
D.Use Principal Component Analysis (PCA) to reduce the feature space to the top 50 principal components that explain 95% of the variance.
AnswerD

PCA transforms the 500 correlated features into orthogonal components ranked by explained variance, so 50 components capture 95% of the variance. This satisfies the stem's dual constraint: cut dimensionality while preserving variance, unlike feature selection, which discards rather than recombines information.

Why this answer

Principal Component Analysis (PCA) is a linear dimensionality reduction technique that transforms the original high-dimensional feature space into a set of orthogonal principal components, ordered by the amount of variance they capture. By selecting the top 50 components that explain 95% of the variance, the team effectively reduces the feature set from over 500 dimensions while preserving the most informative structure in the data, directly addressing the goals of lowering computational cost and mitigating overfitting.

Exam trap

CompTIA AI+ exam questions often test the distinction between dimensionality reduction techniques (PCA) and feature selection methods (Lasso, correlation-based selection) or visualization tools (t-SNE), expecting candidates to recognize that PCA is the only option that explicitly reduces dimensionality while preserving maximum variance in a way that is suitable for downstream modeling.

How to eliminate wrong answers

Option A is wrong because L1 regularization (Lasso) is a feature selection method applied during model training, not a dimensionality reduction technique applied to the feature set before modeling; it does not reduce the number of features in the dataset itself and can still leave high-dimensional data for preprocessing. Option B is wrong because t-SNE is a non-linear visualization technique primarily used for exploring high-dimensional data in 2D or 3D plots; it does not preserve global variance structure, is non-deterministic, and cannot be applied to new unseen data points, making it unsuitable for preprocessing in a production recommendation system. Option C is wrong because selecting only features with high correlation to the target variable ignores interactions between features and can discard features that, while individually weakly correlated, contribute significantly to variance when combined; this approach risks losing valuable information and is not a principled variance-preserving dimensionality reduction method.

63
MCQeasy

Refer to the exhibit. A data engineer is training a binary classification neural network. The loss fluctuates and does not converge. Which hyperparameter adjustment is most likely to stabilize training?

A.Change the activation to tanh
B.Add dropout after each layer
C.Decrease the learning rate
D.Increase the number of units in the first dense layer
AnswerC

Lowering the learning rate reduces the size of each gradient descent step, preventing the optimiser from overshooting minima that cause the loss to oscillate rather than converge. For a binary classification network with fluctuating, non-converging loss, this directly addresses the instability constraint in the stem.

Why this answer

Fluctuating loss that fails to converge during neural network training is a classic sign of an excessively high learning rate, causing the optimizer to overshoot the minimum. Decreasing the learning rate allows the gradient descent updates to take smaller, more stable steps, which smooths the loss curve and promotes convergence.

Exam trap

The CompTIA AI exam often tests the misconception that regularization techniques like dropout or activation changes can fix convergence issues, when in fact the most direct hyperparameter for stabilizing training loss is the learning rate.

How to eliminate wrong answers

Option A is wrong because changing the activation to tanh does not directly address the stability of gradient updates; tanh can help with vanishing gradients in deep networks but does not fix loss oscillation caused by a high learning rate. Option B is wrong because adding dropout is a regularization technique that reduces overfitting by randomly dropping neurons, but it does not stabilize the loss curve during training and may even increase variance in the loss. Option D is wrong because increasing the number of units in the first dense layer increases model capacity and can lead to more complex loss landscapes, potentially exacerbating instability rather than stabilizing training.

64
Multi-Selectmedium

A retail analytics team is preparing a dataset for a demand forecasting model. The dataset contains a 'store_id' column with several thousand unique values, a 'product_category' column with about twenty values, and a 'day_of_week' column. The team wants to encode these categorical variables so a tree-based model can use them effectively without creating an enormous number of columns. Which TWO encoding approaches are most appropriate? (Choose two.)

Select 2 answers
A.Use target encoding for 'store_id', replacing each store with a statistic such as the mean of the target computed from training folds.
B.Use one-hot encoding for 'store_id' so every store gets its own binary column.
C.Use binary encoding for 'day_of_week' so each day is represented by a binary code across multiple columns.
D.Use one-hot encoding for 'product_category' and 'day_of_week' because they have low cardinality.
E.Use label encoding for 'product_category' so categories become arbitrary integers based on alphabetical order.
AnswersA, D

Target encoding maps each high-cardinality store to a single numeric value derived from the target, so thousands of stores become one column instead of thousands of indicator columns. When computed with out-of-fold statistics it avoids leaking the current row's target into its own feature, and tree models can split on the resulting numeric value efficiently.

Why this answer

High-cardinality identifiers like store_id need a compact numeric representation, and out-of-fold target encoding provides one column that captures store-level signal without exploding dimensionality. Low-cardinality nominal features like product_category and day_of_week are best handled with one-hot encoding, which avoids implying an order and keeps the feature space small. Together these choices match encoding strategy to cardinality.

Exam trap

The trap here is applying one encoding method uniformly across all categorical features instead of matching the technique to each feature's cardinality.

65
MCQmedium

Refer to the exhibit. A data engineer runs a validation report on the customers table. The "income" column has 12 null values. Which imputation strategy is most appropriate for this column?

A.Remove rows with null income
B.Replace nulls with the median income per region
C.Replace nulls with 0
D.Replace nulls with the mean income of the entire dataset
AnswerB

Income is typically right-skewed and contains outliers, so the mean would be distorted; the median is robust. Computing it per region preserves local wage differences, satisfying the requirement to impute the 12 nulls without flattening genuine geographic variation in the column.

Why this answer

Imputing missing income values with the median per region preserves the central tendency of each regional subgroup, which is robust to outliers and maintains the distributional characteristics of the data. This strategy is particularly appropriate for income data, which often exhibits skewness and regional variation, ensuring that the imputed values are contextually relevant and do not distort downstream analytics or machine learning models.

Exam trap

CompTIA often tests the misconception that a global mean or median is always the safest imputation, when in fact ignoring subgroup structure (like region) can introduce significant bias and violate the assumption of missing-at-random conditioned on observed features.

How to eliminate wrong answers

Option A is wrong because removing rows with null income reduces the dataset size and can introduce bias if the missingness is not completely random, potentially degrading model performance and statistical power. Option C is wrong because replacing nulls with 0 is arbitrary and unrealistic for income data, introducing a strong downward bias that can severely skew summary statistics and mislead any analysis or model training. Option D is wrong because replacing nulls with the mean income of the entire dataset ignores regional heterogeneity and is sensitive to outliers, which can inflate variance and produce imputed values that are not representative of the local income distribution.

66
MCQmedium

A data engineer is building a pipeline to ingest streaming data from IoT sensors. Which data storage solution is best suited for real-time analytics on timestamped sensor readings?

A.Data warehouse
B.Relational database
C.Data lake
D.Time-series database
AnswerD

Time-series databases index on timestamp and exploit temporal locality, giving fast range scans and downsampling across sensor readings. This directly satisfies the real-time analytics constraint on timestamped IoT data, where relational stores would bottleneck on high-velocity writes and time-window aggregations.

Why this answer

Time-series databases (TSDBs) are optimized for high-ingest rates of timestamped data and provide efficient downsampling, retention policies, and time-based aggregation functions. For IoT sensor streaming, a TSDB like InfluxDB or TimescaleDB delivers sub-second query performance on time-range scans, which is essential for real-time analytics.

Exam trap

CompTIA often tests the misconception that 'any database can handle time-series data if you add a timestamp column,' ignoring the fundamental architectural differences in storage engines, indexing, and write optimization that make TSDBs the only viable choice for real-time streaming analytics.

How to eliminate wrong answers

Option A is wrong because data warehouses (e.g., Snowflake, Redshift) are designed for batch-oriented, structured querying of historical data and cannot sustain the high write throughput or low-latency time-range scans required for streaming sensor data. Option B is wrong because relational databases (e.g., PostgreSQL, MySQL) use row-based storage and B-tree indexes that degrade under continuous time-series inserts, leading to write contention and slow time-range queries. Option C is wrong because data lakes (e.g., S3, ADLS) store raw data in object storage with no indexing or time-ordering, making real-time analytics impossible due to high read latency and lack of native time-series functions.

67
MCQeasy

A team is using a pre-trained language model for sentiment analysis. They want to adapt it to a specific domain with limited labeled data. Which approach is most efficient?

A.Fine-tune the pre-trained model on domain data
B.Use the pre-trained model as is
C.Train a new model from scratch
D.Ensemble multiple pre-trained models
AnswerA

Fine-tuning updates the pre-trained weights on the small domain-labelled set, adapting the model's representations to domain-specific vocabulary and sentiment patterns. This leverages the existing general language knowledge, so far fewer labelled examples are needed than training from scratch.

Why this answer

Fine-tuning a pre-trained language model on domain-specific labeled data is the most efficient approach because it leverages the general language understanding learned from large corpora while adapting to the target domain with minimal additional data. This process uses transfer learning, where only the final layers or a subset of parameters are updated, significantly reducing the amount of labeled data and compute required compared to training from scratch.

Exam trap

The AI0-001 exam often tests the misconception that a pre-trained model can be used directly for any domain without adaptation, leading candidates to choose Option B, but the trap here is that domain-specific tasks require fine-tuning to align the model's representations with the target data distribution.

How to eliminate wrong answers

Option B is wrong because using the pre-trained model as-is (zero-shot inference) typically yields poor performance on domain-specific sentiment analysis due to vocabulary and context mismatches, as the model was not exposed to domain-specific jargon or sentiment nuances. Option C is wrong because training a new model from scratch requires a massive labeled dataset (often millions of examples) and extensive computational resources, which contradicts the constraint of limited labeled data. Option D is wrong because ensembling multiple pre-trained models without fine-tuning them on domain data does not address the domain adaptation problem; it merely averages their general predictions, which may still be inaccurate for domain-specific sentiment.

Ready to test yourself?

Try a timed practice session using only Ai Models Data Engineering questions.