Courseiva

CCNA Aio Ai Infrastructure Questions

75 of 114 questions · Page 1/2 · Aio Ai Infrastructure topic · Answers revealed

1
MCQeasy

A developer is using Hugging Face Transformers to fine-tune a BERT model for sentiment analysis. They want to track experiments, log metrics, and compare runs. Which MLOps tool should they integrate?

A.Apache Airflow
B.Docker
C.Kubeflow
D.MLflow
AnswerD

MLflow provides experiment tracking, metric logging and run comparison, integrating directly with Hugging Face Transformers training loops. It satisfies the stated need to track experiments and compare runs, which raw training scripts alone cannot deliver.

Why this answer

MLflow is the correct choice because it is purpose-built for experiment tracking, metric logging, and run comparison in machine learning workflows. It provides an API to log parameters, metrics, and artifacts, and its UI allows easy comparison of different fine-tuning runs, which directly matches the developer's need to track experiments and compare runs for a BERT sentiment analysis model.

Exam trap

CompTIA often tests the distinction between infrastructure tools (Airflow, Docker, Kubeflow) and ML-specific experiment tracking tools (MLflow), trapping candidates who confuse orchestration or containerization with MLOps tracking capabilities.

How to eliminate wrong answers

Option A is wrong because Apache Airflow is a workflow orchestration tool for scheduling and managing DAGs (Directed Acyclic Graphs) of tasks, not for experiment tracking or metric logging; it lacks native ML run comparison capabilities. Option B is wrong because Docker is a containerization platform for packaging applications and dependencies, not an MLOps tool for logging metrics or comparing experiments; it provides environment consistency but no tracking or logging features. Option C is wrong because Kubeflow is a Kubernetes-native platform for deploying and managing ML pipelines at scale, but it is overkill for simple experiment tracking and does not offer the lightweight, focused metric logging and run comparison that MLflow provides out of the box.

2
MCQhard

An ML team deploys a model on edge devices using INT8 quantization. They notice a significant drop in accuracy on a subset of classes. Which technique should they apply to recover accuracy without increasing model size?

A.Use pruning to remove less important weights
B.Increase the model architecture size
C.Switch to FP16 quantization
D.Apply quantization-aware training (QAT)
AnswerD

Quantization-aware training simulates INT8 rounding during the forward pass while keeping weights in higher precision for gradient updates, letting the model learn to compensate for that error. This recovers accuracy on the affected classes while the deployed artefact stays INT8, so model size is unchanged.

Why this answer

Quantization-aware training (QAT) simulates INT8 quantization effects during the forward pass of training, allowing the model to learn weights and activations that are more robust to the lower precision. This recovers accuracy lost during post-training quantization without increasing the model's size, as the architecture and number of parameters remain unchanged.

Exam trap

CompTIA AI often tests the misconception that post-training quantization is always lossless, leading candidates to overlook the need for QAT when accuracy drops on specific classes due to uneven weight distributions.

How to eliminate wrong answers

Option A is wrong because pruning reduces model size by removing less important weights, which does not directly address the accuracy drop caused by INT8 quantization and may further degrade performance. Option B is wrong because increasing the model architecture size would increase the model's memory footprint and latency, contradicting the requirement to not increase model size. Option C is wrong because switching to FP16 quantization uses 16-bit floating point, which increases the model size compared to INT8 and does not meet the constraint of maintaining the same model size.

3
MCQeasy

A company wants to build a real-time anomaly detection system for IoT sensor data using edge AI. The model must run on resource-constrained devices with minimal power consumption. Which model optimization technique is MOST important?

A.Use FP32 precision
B.Model quantization (INT8)
C.Increase the number of layers
D.Use a larger batch size
AnswerB

INT8 quantization shrinks weights and activations to 8-bit integers, cutting memory footprint and enabling integer arithmetic that draws far less power than FP32. This directly satisfies the stem's constraint of running on resource-constrained edge devices with minimal power consumption.

Why this answer

Quantization reduces model precision (e.g., FP32 to INT8), decreasing model size and computation, which is critical for resource-constrained edge devices.

4
Multi-Selectmedium

A startup is building a retrieval-augmented generation (RAG) application that must answer questions over a 500,000-document internal knowledge base with low query latency. They plan to use a vector database. Which TWO design choices best support fast, scalable similarity search? (Choose two.)

Select 2 answers
A.Store the full text of each document inside the vector database and run keyword search alongside vector search.
B.Choose embeddings with a smaller dimensionality that still preserve semantic quality for the domain.
C.Recompute embeddings for all documents on every query to ensure freshness.
D.Use an approximate nearest neighbor (ANN) index such as HNSW instead of exact brute-force search.
E.Disable metadata filtering so the index can scan all vectors uniformly.
AnswersB, D

Vector search cost scales with dimensionality because distance computations and index memory grow accordingly. Selecting a lower-dimensional embedding that retains domain-relevant semantics reduces memory footprint and speeds up both index construction and query time. This is a legitimate way to support fast, scalable similarity search without abandoning semantic fidelity.

Why this answer

ANN indexing with HNSW and lower-dimensional embeddings both reduce the computational cost of similarity search on a large corpus. HNSW avoids exhaustive comparison, while reduced dimensionality shrinks distance calculations and memory use. Together they enable low-latency retrieval at the 500,000-document scale without sacrificing meaningful semantic accuracy.

Exam trap

The trap here is assuming that adding more features, such as keyword search or full-text storage, automatically improves performance, when the question is about latency and scalability of similarity search.

5
MCQeasy

A data scientist is choosing a hardware accelerator for training a large transformer model. Which of the following is specifically designed for deep learning workloads and offers the highest throughput for matrix multiplications?

A.TPU
B.GPU
C.NPU
D.CPU
AnswerA

TPUs are Google's custom ASICs built specifically for tensor computations, delivering the highest throughput for matrix multiplications in deep learning.

Why this answer

The TPU (Tensor Processing Unit) is an application-specific integrated circuit (ASIC) designed by Google specifically to accelerate deep learning workloads. Its systolic array architecture is optimized for the matrix multiplications and convolutions that dominate transformer model training, delivering the highest throughput among the listed options for these operations.

Exam trap

This question tests the distinction between hardware designed for training versus inference. The trap is that candidates may choose GPU because it is the most common deep learning accelerator, overlooking that TPU is purpose-built for the highest matrix multiplication throughput in training workloads.

How to eliminate wrong answers

Option B (GPU) is wrong because while GPUs are widely used for deep learning and offer high parallelism, they are general-purpose processors originally designed for graphics rendering, not specifically optimized for the dense matrix operations in transformer training. Option C (NPU) is wrong because Neural Processing Units are typically designed for low-power inference on edge devices, not for high-throughput training of large models. Option D (CPU) is wrong because CPUs are general-purpose processors optimized for sequential tasks and low-latency operations, lacking the massive parallel compute units and specialized matrix multiplication hardware needed for efficient transformer training.

6
MCQeasy

A retail company wants to add natural language search to its product catalog. The team plans to convert product descriptions and customer queries into embeddings so that semantically similar items surface even when the wording differs. They need an embedding model that maps text into a dense vector space where cosine similarity reflects meaning. Which type of model should they use?

A.An automatic speech recognition model
B.A named entity recognition model
C.A text embedding model
D.A sentiment analysis model
AnswerC

Text embedding models are trained so that semantically similar sentences land close together in a dense vector space, which is exactly what semantic product search needs. Encoding both catalog descriptions and incoming queries with the same model lets cosine similarity rank items by meaning rather than exact keyword overlap, so a query about a waterproof jacket can match related listings.

Why this answer

Semantic search depends on representing both queries and documents as vectors whose proximity encodes meaning. A text embedding model is trained precisely for that, producing dense representations where cosine similarity correlates with semantic relatedness. Using one model for both catalog text and queries keeps them in the same space, enabling relevant matches despite different phrasing.

Exam trap

The trap here is confusing any NLP model with an embedding model, when only models trained for representation learning yield the dense vectors similarity search requires.

7
MCQeasy

An ML engineer wants to deploy a model as a REST API that can scale to handle thousands of inference requests per second. Which serving approach is most appropriate?

A.Export the model to ONNX format and use a batch processing pipeline
B.Use gRPC streaming for all inference requests
C.Run the model directly on the client device
D.Deploy the model as a REST API endpoint using a containerized inference server
AnswerD

A containerised inference server exposes the model as a REST endpoint and supports horizontal scaling across replicas behind a load balancer, meeting the thousands-of-requests-per-second requirement. Containers also package dependencies consistently, enabling elastic autoscaling of inference capacity.

Why this answer

Deploying the model as a REST API endpoint using a containerized inference server (e.g., TensorFlow Serving, TorchServe, or NVIDIA Triton Inference Server) is the most appropriate approach for handling thousands of inference requests per second. These servers are designed for high-throughput, low-latency serving, support horizontal scaling via load balancers, and provide built-in batching and model versioning. REST APIs are stateless and can be easily integrated with existing web infrastructure, making them ideal for production-scale inference.

Exam trap

The AI0-001 exam often tests the distinction between serving infrastructure (REST API with containerized server) and data processing pipelines (batch) or communication protocols (gRPC), leading candidates to confuse a transport mechanism or batch method with a scalable serving architecture.

How to eliminate wrong answers

Option A is wrong because exporting to ONNX and using a batch processing pipeline is designed for offline/batch inference, not for real-time REST API serving with thousands of requests per second; batch pipelines introduce latency and are not suitable for synchronous, low-latency inference. Option B is wrong because gRPC streaming is a communication protocol that can be used for inference, but it is not a serving approach itself; moreover, gRPC streaming is typically used for bidirectional or long-lived streams, not for high-volume stateless REST API requests, and it adds complexity without inherent scalability benefits over REST for this use case. Option C is wrong because running the model directly on the client device (edge inference) offloads computation from the server but does not provide a centralized REST API; it also introduces challenges with model updates, device heterogeneity, and security, and is not a server-side serving approach.

8
MCQmedium

A machine learning team is training a large transformer model on a text corpus. They need to reduce training time while maintaining model accuracy. Which hardware configuration would be MOST effective for this task?

A.Use a high-core-count CPU with large RAM
B.Use a cluster of GPUs with data parallelism
C.Use a single GPU with model parallelism
D.Use a single TPU with model parallelism
AnswerB

Data parallelism distributes each batch across many GPUs, each holding a full model replica and synchronising gradients, which cuts wall-clock training time substantially. This scales effectively for large transformer models while preserving accuracy through equivalent gradient updates.

Why this answer

Training large transformer models is computationally intensive, and data parallelism across a cluster of GPUs allows the model to process multiple batches simultaneously, dramatically reducing training time. Each GPU holds a full copy of the model and processes a different subset of the data, with gradients synchronized across devices. This approach scales well and maintains accuracy as long as the effective batch size is tuned appropriately.

Exam trap

AI0-001 often tests the misconception that a single powerful device (TPU or GPU) with model parallelism is better than a multi-GPU cluster; candidates may pick option D because TPUs are marketed for transformers, but the question emphasizes reducing training time, which favors data parallelism across many GPUs.

How to eliminate wrong answers

Option A is wrong because high-core-count CPUs with large RAM are far slower than GPUs for the matrix multiplications and attention operations in transformers; CPUs lack the massive parallelism and high memory bandwidth needed for efficient deep learning. Option C is wrong because a single GPU with model parallelism splits the model across layers within one device, which does not reduce training time as effectively as scaling across multiple GPUs; it also introduces communication overhead and is limited by the single GPU's memory and compute. Option D is wrong because a single TPU with model parallelism, while powerful, is less effective than a cluster of GPUs for reducing training time on large transformers; TPUs are optimized for specific workloads and may require code changes, and a single device cannot match the throughput of a multi-GPU cluster.

9
Multi-Selectmedium

A team is selecting a vector database for a RAG application that requires low-latency similarity search on millions of embeddings. They prioritize ease of use and fully managed cloud service. Which TWO options meet these requirements?

Select 2 answers
A.Pinecone
B.pgvector
C.Chroma
D.Weaviate
E.Milvus
AnswersA, D

Pinecone is a fully managed, cloud-native vector database with low-latency similarity search, ideal for production RAG.

Why this answer

Pinecone is a fully managed vector database designed for production-scale RAG applications, offering low-latency similarity search on millions of embeddings without requiring users to manage infrastructure. Its serverless architecture and simple API align directly with the team's priorities of ease of use and a fully managed cloud service. Weaviate also offers a fully managed cloud service (Weaviate Cloud) with low-latency similarity search and an easy-to-use API, making it a second valid choice.

By contrast, pgvector is a PostgreSQL extension that is not itself a fully managed cloud service; Chroma is primarily an open-source embedded/local database; and Milvus is typically self-managed (Zilliz Cloud is a separate managed offering), so these do not inherently satisfy the 'fully managed cloud service' requirement.

Exam trap

CompTIA often tests the distinction between open-source, self-managed tools and fully managed cloud services, where candidates may incorrectly assume that any popular vector database (like Milvus or pgvector) inherently provides a managed cloud experience without checking the deployment model.

10
MCQhard

A healthcare AI team is training a model to predict patient readmission risk from electronic health records. The dataset contains sensitive patient data and must comply with HIPAA. They need to ensure that the model training process does not expose protected health information (PHI) and that the model does not memorize individual patient data. Which technique should they implement?

A.Data augmentation with synthetic records
B.Federated learning across hospitals
C.Homomorphic encryption of the model weights
D.Differential privacy during training
AnswerD

Differential privacy adds controlled noise to the training process, ensuring that the inclusion or exclusion of any single patient's data does not significantly affect the model's output. This prevents memorization of individual records and helps comply with HIPAA by protecting PHI. It is a rigorous mathematical framework for privacy-preserving machine learning.

Why this answer

Differential privacy is the only technique listed that provides a formal guarantee that the model does not memorize individual patient data. By injecting noise during training, it ensures that the model's predictions are statistically similar whether or not any single patient's data is included. This directly addresses the HIPAA compliance and privacy concerns in the scenario.

Exam trap

The trap here is confusing privacy-preserving techniques that protect data in transit or storage with those that prevent model memorization during training.

11
Multi-Selecthard

A company is building a multi-modal AI application that processes text, images, and audio. They need a unified platform to store embeddings for all modalities, perform hybrid search (vector + metadata filtering), and scale to millions of vectors. Which THREE services are suitable for this purpose? (Choose THREE.)

Select 3 answers
A.Weaviate
B.Amazon S3
C.Snowflake
D.pgvector (PostgreSQL extension)
E.Pinecone
AnswersA, D, E

Weaviate is a vector database with hybrid search and multi-modal support.

Why this answer

Weaviate is a purpose-built vector database that natively supports multi-modal embeddings (text, images, audio) through its vectorizer modules and provides hybrid search combining vector similarity with metadata filtering (e.g., using GraphQL or REST APIs). It is designed to scale to millions of vectors with built-in sharding and replication, making it a strong fit for the described unified platform.

Exam trap

CompTIA often tests the distinction between general-purpose storage (S3) or analytics platforms (Snowflake) and purpose-built vector databases, leading candidates to mistakenly choose services that store data but lack native vector search and hybrid filtering capabilities.

12
MCQhard

A team wants to deploy a large language model on edge devices with limited memory and compute. They need to reduce model size by at least 50% while preserving accuracy. Which combination of techniques is most effective?

A.Apply INT8 quantization and weight pruning
B.Distill the model into a smaller architecture without quantization or pruning
C.Use FP32 precision and increase batch size
D.Use FP16 quantization and add more layers
AnswerA

INT8 quantization shrinks weights from 32-bit to 8-bit, cutting size roughly 75%, while weight pruning removes redundant connections. Combined, they exceed the 50% reduction target with minimal accuracy loss, fitting edge memory and compute limits.

Why this answer

INT8 quantization reduces the precision of weights and activations from 32-bit to 8-bit, cutting memory usage by approximately 75% for those tensors, while weight pruning removes redundant connections, often achieving over 50% size reduction with minimal accuracy loss when combined. Together, they directly address the constraints of edge devices by shrinking the model footprint and computational requirements without requiring a complete architecture redesign.

Exam trap

The exam often tests the misconception that a single technique (like distillation or FP16) is sufficient for aggressive size reduction, when in reality, combining complementary compression methods (quantization and pruning) is necessary to meet both the 50% size reduction and accuracy preservation requirements on edge devices.

How to eliminate wrong answers

Option B is wrong because knowledge distillation alone reduces model size by training a smaller student network, but without quantization or pruning, the student model may still exceed the 50% reduction target or suffer significant accuracy loss if the architecture is not aggressively compressed. Option C is wrong because using FP32 precision and increasing batch size actually increases memory and compute demands, making it unsuitable for resource-constrained edge devices. Option D is wrong because FP16 quantization provides only a 50% memory reduction (not guaranteed to meet the target when combined with adding layers, which increases model size and complexity, often negating the quantization benefit and degrading accuracy on edge hardware without native FP16 support.

13
MCQmedium

A data science team is deploying a deep learning model for real-time inference on edge devices with limited power and memory. Which model optimisation technique would be MOST effective for reducing latency and memory footprint while maintaining acceptable accuracy?

A.Use a larger batch size during inference
B.Train the model for more epochs to improve convergence
C.Apply quantisation to convert weights from FP32 to INT8
D.Increase the number of layers to improve feature extraction
AnswerC

Quantisation converts FP32 weights to INT8, cutting memory footprint roughly fourfold and enabling integer arithmetic that executes faster on edge hardware, directly satisfying the limited power and memory constraint. Accuracy loss stays acceptable because scaling factors preserve the weight distribution, making it ideal for real-time inference on constrained devices.

Why this answer

Quantization reduces the precision of model weights from 32-bit floating point (FP32) to 8-bit integer (INT8), which directly cuts memory usage by 75% and accelerates inference on edge devices by leveraging integer arithmetic. This technique is specifically designed for resource-constrained environments where power and memory are limited, and it typically preserves accuracy within 1-2% of the original model.

Exam trap

CompTIA often tests the misconception that increasing model complexity (more layers or epochs) improves deployment performance, when in fact the opposite is true for edge inference; candidates may confuse training optimization with inference optimization.

How to eliminate wrong answers

Option A is wrong because increasing batch size during inference increases memory consumption and latency on edge devices, as it requires processing multiple inputs simultaneously, which is counterproductive for real-time, low-latency requirements. Option B is wrong because training for more epochs improves convergence and accuracy but does not reduce model size or inference latency; it may even lead to overfitting without any benefit to deployment efficiency. Option D is wrong because adding more layers increases the model's parameter count, memory footprint, and computational latency, directly opposing the goal of reducing resource usage on edge devices.

14
MCQmedium

A data scientist needs to deploy a PyTorch model to production with low-latency inference. The model must be served as a REST API and should support GPU acceleration. Which combination of tools is MOST suitable for this task?

A.ONNX runtime with a gRPC endpoint on a CPU-only node
B.Apache Spark with MLlib to serve the model in batch mode
C.Docker container with a FastAPI application and Nvidia GPU support
D.Kubeflow Pipelines to deploy the model as a scheduled job
AnswerC

Docker packages the model and FastAPI exposes it as a REST API, while Nvidia GPU support enables CUDA acceleration for low-latency inference. This combination satisfies both the REST API and GPU acceleration constraints for the PyTorch model.

Why this answer

It combines Docker containerization with FastAPI for a lightweight REST API and NVIDIA GPU support (via nvidia-docker or NVIDIA Container Toolkit) to enable low-latency GPU-accelerated inference. This stack directly meets the requirements of low-latency inference, REST API serving, and GPU acceleration without unnecessary overhead.

Exam trap

CompTIA often tests the distinction between batch/offline processing tools (like Spark or Kubeflow Pipelines) and real-time serving frameworks, leading candidates to confuse orchestration or batch tools with low-latency inference solutions.

How to eliminate wrong answers

Option A is wrong because ONNX Runtime with a gRPC endpoint on a CPU-only node cannot provide GPU acceleration, which is explicitly required. Option B is wrong because Apache Spark with MLlib is designed for distributed batch processing and large-scale data pipelines, not for low-latency real-time REST API serving of a single PyTorch model. Option D is wrong because Kubeflow Pipelines is a workflow orchestration tool for scheduling and managing ML pipelines, not a real-time inference serving solution; it lacks native REST API endpoints for low-latency inference.

15
MCQhard

A company uses Azure OpenAI to generate customer support responses. The team notices that repeated queries with similar context incur high costs due to token usage. They want to reduce costs without affecting response quality. Which strategy is MOST effective?

A.Use a larger model to improve efficiency
B.Increase the frequency penalty
C.Reduce the max_tokens parameter
D.Implement prompt caching
AnswerD

Prompt caching stores and reuses the computed key-value representations of repeated context, so identical or similar prefixes are billed at a reduced cached-token rate rather than full input tokens. This directly cuts token costs on recurring queries while returning the same model output, preserving response quality.

Why this answer

Prompt caching stores and reuses tokens from previous queries, reducing token consumption for similar requests and lowering costs without quality loss.

16
MCQeasy

An AI developer needs to store large amounts of unstructured data (e.g., images, logs) for training datasets. Which cloud storage solution is purpose-built for data lakes?

A.Amazon DynamoDB
B.Amazon RDS
C.Amazon S3
D.Amazon Redshift
AnswerC

Amazon S3 provides object storage with flat namespace, unlimited scalability and high durability, purpose-built for data lakes holding unstructured images and logs. It satisfies the stem's requirement for storing large volumes of unstructured training data, unlike block or file storage. S3's decoupling of metadata from objects suits analytics engines querying datasets directly.

Why this answer

Amazon S3 is purpose-built for data lakes because it provides virtually unlimited scalability, high durability (99.999999999% or 11 nines), and supports any type of unstructured data (images, logs, videos) with a flat object storage architecture. Its integration with AWS Glue, Athena, and Lake Formation enables schema-on-read analytics, making it the foundational service for building a data lake on AWS.

Exam trap

Candidates often mistake Redshift for a data lake solution because they associate 'data' with 'warehouse' rather than recognizing that data lakes require raw object storage.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for low-latency transactional workloads, not for storing large volumes of unstructured data for analytics. Option B is wrong because Amazon RDS is a relational database service for structured data with fixed schemas, and it cannot scale to petabyte-scale unstructured data storage. Option D is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for structured, columnar data and SQL-based analytics, not for storing raw unstructured data like images or logs.

17
MCQmedium

A data scientist is building a recommendation system using Apache Spark for feature engineering. They need to process streaming user click data in real-time before feeding into the model. Which tool should they use for the streaming data ingestion?

A.Amazon S3
B.Apache Kafka
C.Airflow
D.Snowflake
AnswerB

Kafka provides a distributed, partitioned, replicated commit log that ingests high-throughput click streams with low latency and durable buffering, decoupling producers from Spark Structured Streaming consumers. This satisfies the requirement to process real-time streaming click data before feature engineering.

Why this answer

Apache Kafka is the correct choice because it is a distributed streaming platform designed for high-throughput, fault-tolerant, real-time data ingestion. It acts as a durable message broker that can ingest streaming click data and make it available for Spark Structured Streaming to process in micro-batches or continuous processing mode, which is essential for real-time feature engineering in a recommendation system.

Exam trap

CompTIA often tests the distinction between storage, orchestration, and streaming tools, and the trap here is that candidates confuse batch-oriented tools like S3 or Airflow with real-time streaming ingestion, overlooking Kafka's role as a dedicated event streaming platform.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is an object storage service, not a streaming ingestion tool; it lacks the low-latency, pub-sub messaging capabilities required for real-time data streaming. Option C is wrong because Airflow is a workflow orchestration tool for scheduling batch jobs, not a real-time streaming ingestion platform; it cannot handle continuous, event-driven data streams. Option D is wrong because Snowflake is a cloud-based data warehouse optimized for analytical queries on structured data, not for real-time streaming ingestion; it does not provide a pub-sub or message queue interface for live click data.

18
MCQhard

A cybersecurity firm is building an anomaly detection system for network traffic. The dataset contains millions of connection records with dozens of features, but only 0.1% are labeled as malicious. The team needs a model that can flag suspicious connections while minimizing false positives that overwhelm analysts. Which approach is most appropriate?

A.Apply principal component analysis (PCA) to reduce dimensionality, then use a threshold on reconstruction error to detect anomalies.
B.Use k-means clustering with k set to the number of known attack types and treat points far from centroids as anomalies.
C.Train an isolation forest or autoencoder on the normal traffic to learn its structure, then flag deviations as anomalies.
D.Train a supervised gradient boosting classifier on the labeled data and use class weights to handle imbalance.
AnswerC

With only 0.1% malicious labels, supervised classification is impractical due to extreme class imbalance. Unsupervised anomaly detection methods like isolation forest and autoencoders learn the distribution of normal traffic and flag deviations, requiring no labels. This directly suits the scenario, and by tuning the contamination parameter or reconstruction error threshold, the team can control the false positive rate to keep analysts from being overwhelmed.

Why this answer

Extreme class imbalance with only 0.1% malicious labels makes supervised learning unreliable. Unsupervised anomaly detection methods like isolation forest and autoencoders learn the structure of normal traffic without requiring labels, then flag deviations. Isolation forest isolates anomalies by random partitioning, while autoencoders flag high reconstruction error.

Both allow threshold tuning to balance detection and false positives, which is critical for analyst workload. This makes them the most appropriate choice for the cybersecurity scenario.

Exam trap

The trap here is assuming that class weighting or dimensionality reduction alone can overcome a 0.1% positive rate, when the scarcity of labels makes unsupervised anomaly detection the only practical path.

19
MCQhard

A team is deploying a model on Kubernetes using Kubeflow. They want to automatically scale the number of inference pods based on request latency. Which Kubernetes-native feature should they configure?

A.Horizontal Pod Autoscaler (HPA) with custom metrics
B.Kubeflow Pipelines component
C.Cluster Autoscaler
D.Vertical Pod Autoscaler (VPA)
AnswerA

HPA scales pod replicas from metrics, and custom metrics let it target request latency rather than CPU. This satisfies the latency-based scaling constraint, since the default resource metrics cannot express latency and would not react to slow inference responses.

Why this answer

The Horizontal Pod Autoscaler (HPA) with custom metrics is the correct choice because it allows scaling based on application-level metrics like request latency, not just CPU or memory. By configuring HPA to use a custom metric (e.g., from Prometheus or a metrics adapter), the team can automatically adjust the number of inference pods to maintain target latency thresholds, which is essential for responsive inference serving.

Exam trap

The distinction between pod-level scaling (HPA) and node-level scaling (Cluster Autoscaler) is important. The trap is that candidates may confuse Cluster Autoscaler with pod autoscaling, or assume VPA can handle latency-based scaling when it only adjusts resource limits.

How to eliminate wrong answers

Option B is wrong because Kubeflow Pipelines is a workflow orchestration component for building and managing ML pipelines, not a scaling mechanism; it cannot directly scale pods based on latency. Option C is wrong because Cluster Autoscaler adjusts the number of nodes in the Kubernetes cluster, not the number of pods, and does not respond to request latency metrics. Option D is wrong because Vertical Pod Autoscaler (VPA) adjusts CPU/memory resource requests for existing pods, not the number of pods, and is not designed for latency-based scaling.

20
Multi-Selectmedium

A company is building an AI-powered document processing system that extracts information from scanned PDFs. The system must handle varying document layouts and languages. The team wants to use a pre-trained model and fine-tune it on their own data. Which TWO techniques are most appropriate to improve the model's ability to generalize to new document layouts? (Choose two.)

Select 2 answers
A.Data augmentation with random rotations, scaling, and cropping of document images.
B.Reducing the model size by pruning 50% of the weights before fine-tuning.
C.Incorporating a layout-aware pre-training objective such as masked visual-language modeling.
D.Fine-tuning all layers of the model with a very low learning rate.
E.Using a larger batch size during training to stabilize gradients.
AnswersA, C

Data augmentation introduces variability in the training data, simulating different layouts and scanning conditions. This helps the model learn invariant features and improves generalization to unseen document formats. For document processing, augmentations like rotation and scaling are effective because they mimic real-world distortions without requiring new labeled data.

Why this answer

Data augmentation with geometric transformations exposes the model to layout variations, while layout-aware pre-training objectives help the model learn structural relationships. Together, they enhance generalization to unseen document formats. Other options focus on training efficiency or model compression, which do not directly address layout variability.

Exam trap

The trap here is confusing techniques that improve training stability or speed with those that improve generalization to new layouts.

21
Multi-Selectmedium

An organization is building a recommendation system that requires low-latency vector similarity search. They need to store and query millions of embeddings. Which THREE technologies are appropriate for this task?

Select 3 answers
A.Snowflake
B.Amazon S3
C.Weaviate
D.pgvector
E.Pinecone
AnswersC, D, E

Weaviate is a purpose-built vector database that indexes embeddings using HNSW graphs, delivering the low-latency approximate nearest-neighbour similarity search the scenario demands across millions of vectors. It satisfies the scale and latency constraints directly, unlike general-purpose stores lacking native vector indexing.

Why this answer

Weaviate (C) is a purpose-built vector database that indexes embeddings and supports low-latency approximate nearest neighbor (ANN) similarity search over millions of vectors, making it ideal for recommendation systems. pgvector (D) extends PostgreSQL with vector data types and ANN indexes (e.g., HNSW, IVFFlat) so embeddings can be stored and queried with low latency alongside relational data. Pinecone (E) is a fully managed vector database designed specifically for high-performance similarity search at scale, directly matching the low-latency embedding query requirement. Snowflake (A) is a cloud data warehouse optimized for analytical SQL workloads, not sub-second vector similarity search, and Amazon S3 (B) is object storage that can hold embedding files but provides no native vector indexing or similarity query capability.

Exam trap

The trap is confusing general-purpose data stores (Snowflake, S3) with vector databases; candidates may think any storage can handle embeddings, but only specialized vector databases provide the necessary indexing and low-latency search.

22
MCQmedium

A financial institution requires that all AI model predictions be explainable and auditable for regulatory compliance. Which model serving approach should be used to meet these requirements?

A.Use gRPC streaming for lower latency
B.Export the model to ONNX format and run on a dedicated inference server
C.Deploy the model on edge devices to avoid centralised logging
D.Deploy the model as a containerised microservice with REST API and log all request/response pairs
AnswerD

Containerised microservice deployment with REST API request/response logging captures each prediction's inputs and outputs, creating the auditable record regulators require. This satisfies the explainability and auditability constraint without relying on opaque batch or embedded serving.

Why this answer

Logging all request/response pairs provides a complete audit trail, which is essential for regulatory compliance in financial institutions. Containerized microservices with REST APIs are stateless and can be easily integrated with centralized logging systems (e.g., ELK stack) to capture every prediction for explainability and review. This approach ensures that model decisions are transparent and can be traced back to specific inputs, satisfying both explainability and auditability requirements.

Exam trap

CompTIA often tests the misconception that performance optimizations (like gRPC or ONNX) inherently solve compliance requirements, when in fact auditability and explainability depend on explicit logging and traceability mechanisms, not just model format or transport protocol.

How to eliminate wrong answers

Option A is wrong because gRPC streaming focuses on low-latency communication, not on logging or auditability; it does not inherently provide request/response capture for compliance. Option B is wrong because exporting to ONNX and running on a dedicated inference server improves portability and performance but does not automatically log predictions or provide an audit trail; additional logging infrastructure would be required. Option C is wrong because deploying on edge devices avoids centralized logging, which directly contradicts the need for auditable records; edge deployments often lack persistent, centralized storage of prediction history, making regulatory review difficult.

23
MCQmedium

A company uses a cloud-based ML platform to train a model and wants to deploy it for real-time inference. They also need to monitor the endpoint for data drift and retrain automatically. Which feature enables this automated retraining pipeline?

A.ML Pipeline Orchestration
B.Model Debugging
C.Data Labeling Service
D.Model Monitoring
AnswerA

ML pipeline orchestration chains training, evaluation, deployment and monitoring steps into an automated workflow, so a drift trigger can restart training and redeploy the endpoint without manual intervention. This satisfies the requirement for automatic retraining of the real-time inference model.

Why this answer

ML Pipeline Orchestration is the correct answer because it provides a fully managed service for creating, automating, and managing end-to-end machine learning workflows. It allows you to define a pipeline that includes steps for monitoring data drift (via a model monitoring service), triggering retraining jobs, and deploying updated models, enabling the automated retraining pipeline described in the question.

Exam trap

The trap here is that candidates often confuse a model monitoring service's detection capability with the full orchestration needed for automated retraining, assuming that monitoring alone can trigger retraining without a pipeline orchestration service.

How to eliminate wrong answers

Option B is wrong because SageMaker Debugger is designed for debugging training jobs by monitoring system metrics, profiling, and detecting anomalies like vanishing gradients, not for orchestrating automated retraining pipelines. Option C is wrong because SageMaker Ground Truth is a data labeling service that creates high-quality training datasets using human annotators, not for automating model retraining or deployment. Option D is wrong because SageMaker Model Monitor only detects data drift and quality issues by analyzing inference data, but it does not include the orchestration logic to automatically trigger retraining or redeployment; that requires a pipeline service like SageMaker Pipelines.

24
Multi-Selecthard

A retail analytics team is choosing a vector database to power semantic search over millions of product descriptions and support retrieval-augmented generation. The team must keep infrastructure costs predictable and needs fast approximate nearest neighbor queries as the index grows. Which TWO characteristics of approximate nearest neighbor indexing should the team evaluate when selecting the vector database? (Choose two.)

Select 2 answers
A.The recall-versus-latency tradeoff controlled by index search parameters
B.Whether the database enforces ACID transactions across all vector writes
C.Whether the database can generate the embeddings from raw text itself
D.The ability to store product descriptions as fixed-width CHAR columns
E.The memory footprint required to hold the index resident for fast queries
AnswersA, E

ANN indexes such as HNSW expose search-time parameters, for example efSearch, that trade recall against query latency. At millions of vectors, raising the parameter improves the chance of returning true nearest neighbors but increases per-query work and cost. The team must measure this curve on its own data to pick a setting that meets both relevance expectations and predictable infrastructure spend.

Why this answer

ANN indexes are evaluated mainly on the recall-versus-latency curve and on how much memory the index needs to stay resident. Those two factors determine whether semantic search returns relevant products quickly and whether the monthly infrastructure bill stays predictable as the catalog grows. The other listed traits concern transaction semantics, column typing, or embedding generation, none of which govern ANN index performance.

Exam trap

The trap here is conflating database conveniences such as ACID guarantees or built-in embedding with the index characteristics that actually control ANN recall, latency, and memory cost.

25
MCQhard

A company is implementing a retrieval-augmented generation (RAG) pipeline using a vector database. They notice that the retrieved documents often lack relevance to the query. Which adjustment would MOST improve retrieval quality?

A.Use a better embedding model fine-tuned on domain-specific data
B.Increase the chunk size of documents
C.Switch from cosine similarity to Euclidean distance
D.Reduce the number of retrieved documents from 5 to 3
AnswerA

Embedding quality determines vector similarity, so a domain-tuned model maps queries and documents into a space where relevant passages score higher. This directly addresses the relevance constraint in the stem, unlike prompt or chunking tweaks.

Why this answer

Retrieval quality in a RAG pipeline is fundamentally determined by the semantic alignment between query embeddings and document embeddings. A domain-specific fine-tuned embedding model captures the unique terminology, context, and relationships within the company's data, producing vector representations that are far more relevant than those from a generic model. This directly improves the similarity search results in the vector database, leading to higher-quality retrieved documents.

Exam trap

A common mistake is to focus on tuning retrieval parameters (chunk size, distance metric, or k-value) rather than improving the embedding model itself, which is the primary driver of semantic relevance in a RAG pipeline.

How to eliminate wrong answers

Option B is wrong because increasing chunk size can reduce granularity and introduce noise, potentially lowering retrieval precision by mixing irrelevant content with relevant passages. Option C is wrong because cosine similarity and Euclidean distance are both valid distance metrics; switching between them does not inherently improve relevance unless the embedding space is normalized, and cosine similarity is typically preferred for high-dimensional semantic embeddings. Option D is wrong because reducing the number of retrieved documents from 5 to 3 may increase precision but at the cost of recall, and does not address the root cause of poor relevance—the quality of the embeddings themselves.

26
MCQmedium

An MLOps team wants to deploy a trained PyTorch model to production with low latency inference. The model must be interoperable across different frameworks and runtimes. Which approach is BEST?

A.Deploy the native PyTorch model using TorchServe
B.Quantize the model to INT8 and deploy as a TensorFlow Lite model
C.Convert the model to TensorFlow SavedModel and deploy using TensorFlow Serving
D.Export the model to ONNX format and deploy using ONNX Runtime
AnswerD

ONNX provides a framework-agnostic serialisation format, so the exported graph runs identically under ONNX Runtime regardless of the original PyTorch training stack. This satisfies the interoperability constraint across frameworks and runtimes while ONNX Runtime's optimised execution kernels deliver the required low-latency inference.

Why this answer

ONNX (Open Neural Network Exchange) provides a standardized, framework-agnostic format that ensures interoperability across different runtimes and hardware accelerators. By exporting the PyTorch model to ONNX and deploying with ONNX Runtime, the team achieves low-latency inference through graph optimizations and hardware-specific execution providers, while avoiding vendor lock-in.

Exam trap

CompTIA often tests the misconception that framework-native serving (TorchServe, TensorFlow Serving) is the best path for low latency, ignoring the explicit requirement for cross-framework interoperability that ONNX uniquely satisfies.

How to eliminate wrong answers

Option A is wrong because deploying a native PyTorch model with TorchServe locks the inference into the PyTorch ecosystem, violating the requirement for interoperability across different frameworks and runtimes. Option B is wrong because quantizing to INT8 and deploying as a TensorFlow Lite model introduces unnecessary precision loss and framework conversion overhead, and TensorFlow Lite is primarily designed for mobile/edge devices, not general production low-latency serving. Option C is wrong because converting to TensorFlow SavedModel and using TensorFlow Serving ties the deployment to the TensorFlow stack, which does not satisfy the interoperability requirement and adds conversion complexity without the broad runtime support that ONNX provides.

27
MCQeasy

A hospital wants to run a patient-triage natural language model entirely inside its own data center because patient records cannot leave the premises. The IT team needs an inference serving component that exposes an HTTP endpoint, supports model versioning, and can be operated without a managed cloud service. Which technology should the team deploy?

A.A Jupyter notebook that loads the model and calls predict()
B.A managed cloud inference API from a public provider
C.A self-hosted model server such as NVIDIA Triton Inference Server
D.An object storage bucket holding the model weights
AnswerC

Triton Inference Server runs inside the hospital's own data center, exposes HTTP and gRPC inference endpoints, and supports multiple model versions with configurable version policies. It also works with models from several frameworks, so the triage model can be served without a managed cloud service. This satisfies data residency, endpoint, and versioning requirements simultaneously.

Why this answer

The hospital needs inference to stay on-premises while still offering an HTTP endpoint and model versioning. A self-hosted model server such as NVIDIA Triton Inference Server provides exactly that: it runs locally, exposes standard inference endpoints, and manages multiple model versions. Managed APIs, notebooks, and raw object storage each fail at least one of the stated constraints.

Exam trap

The trap here is treating model storage or an interactive notebook as a serving layer, when neither exposes a production HTTP inference endpoint with versioning.

28
MCQeasy

A startup is developing a voice assistant that runs on smart speakers with limited processing power and memory. The team wants to use a pre-trained speech recognition model but needs to reduce its size and latency. Which approach is most suitable?

A.Increase the model's precision to FP64 to improve accuracy.
B.Use the pre-trained model as-is and rely on the smart speaker's hardware acceleration.
C.Deploy the pre-trained model on a cloud server and stream audio for processing.
D.Use knowledge distillation to train a smaller student model from the pre-trained model.
AnswerD

Knowledge distillation transfers knowledge from a large teacher model to a smaller student model, reducing size and latency while maintaining accuracy. This is ideal for smart speakers with limited resources. The student model can be optimized for the specific task, making it a suitable approach for deployment on edge devices.

Why this answer

Knowledge distillation creates a smaller, faster model that retains much of the teacher's performance, making it ideal for edge devices with limited compute and memory. Cloud offloading, higher precision, or using the model unchanged do not address the constraints of size and latency on the smart speaker.

Exam trap

The trap here is assuming that hardware acceleration alone can make a large model run efficiently on a constrained device.

29
MCQhard

A data scientist is training a large language model on a custom dataset using PyTorch on AWS. The training is taking too long due to GPU memory constraints. The team wants to use multiple GPUs across instances with minimal code changes. Which AWS service should they use?

A.AWS Elastic Fabric Adapter (EFA)
B.Amazon SageMaker with distributed training libraries
C.AWS Batch with GPU instances
D.AWS ParallelCluster with Slurm
AnswerB

SageMaker distributed training libraries handle data and model parallelism across multiple GPU instances with minimal PyTorch code changes, directly addressing the GPU memory constraint. It satisfies the stem's requirement to scale beyond a single instance without rewriting the training script.

Why this answer

Amazon SageMaker with distributed training libraries provides built-in support for data and model parallelism across multiple GPUs and instances, requiring minimal code changes. It integrates with PyTorch and handles the orchestration of distributed training, making it the best choice for scaling training.

Exam trap

The trap is confusing infrastructure services (EFA, Batch, ParallelCluster) with managed machine learning services; candidates may pick EFA because it sounds like a networking solution for distributed training, but it lacks the high-level libraries.

How to eliminate wrong answers

Option A is wrong because AWS Elastic Fabric Adapter (EFA) is a network interface for high-performance computing, but it does not provide the distributed training libraries or orchestration; it is a lower-level networking solution. Option C is wrong because AWS Batch with GPU instances manages batch jobs but does not offer built-in distributed training libraries or seamless PyTorch integration for multi-GPU/multi-instance training. Option D is wrong because AWS ParallelCluster with Slurm is a cluster management tool that requires significant setup and code changes to implement distributed training, lacking the managed libraries of SageMaker.

30
Multi-Selectmedium

A healthcare startup needs to deploy an AI model for real-time patient monitoring on IoT devices with limited battery and compute. The model must run locally with minimal latency. Which TWO strategies are most appropriate?

Select 2 answers
A.Apply model distillation to create a smaller student model
B.Deploy the model on a cloud server and stream data
C.Use TensorFlow Lite to convert and run the model on the device
D.Quantize the model to INT8 precision
E.Use ONNX Runtime with a GPU backend
AnswersC, D

TensorFlow Lite is optimized for on-device machine learning, providing low-latency inference on resource-constrained devices.

Why this answer

TensorFlow Lite is specifically designed to run TensorFlow models on resource-constrained edge devices like IoT sensors. It optimizes the model for low latency inference by using a specialized interpreter and hardware acceleration delegates (e.g., NNAPI, GPU), enabling real-time patient monitoring without cloud dependency.

Exam trap

A common misconception is that model distillation alone is sufficient for edge deployment, when in fact it must be combined with a framework like TensorFlow Lite and quantization to meet hardware constraints in a Comptia AI context.

31
MCQmedium

An organization is deploying a large language model on-premises for compliance reasons. They need to serve inference requests with low latency. Which architecture should they use?

A.Use a batch processing system like Apache Spark
B.Containerize the model and deploy it on a Kubernetes cluster with autoscaling
C.Use a serverless function like AWS Lambda
D.Deploy the model as a REST API on a single powerful server
AnswerB

Containerising the model on Kubernetes keeps inference inside the organisation's own infrastructure, satisfying the on-premises compliance constraint. Autoscaling horizontally scales replicas to match request load, and local GPU nodes serve requests without network round trips to a public endpoint, delivering the low latency required.

Why this answer

Containerizing the model and deploying it on a Kubernetes cluster with autoscaling is the correct architecture because it provides horizontal scaling, low-latency inference through load-balanced pods, and supports on-premises deployment for compliance. Kubernetes can automatically scale replicas based on CPU/memory utilization or custom metrics (e.g., request queue depth), ensuring consistent response times under varying load.

Exam trap

CompTIA often tests the misconception that a single powerful server is sufficient for low-latency inference, but the trap is that it ignores the need for horizontal scalability and fault tolerance, which are critical for production workloads.

How to eliminate wrong answers

Option A is wrong because batch processing systems like Apache Spark are designed for large-scale data processing jobs, not real-time inference; they introduce high latency due to job scheduling and data shuffling, making them unsuitable for serving low-latency requests. Option C is wrong because serverless functions like AWS Lambda are typically cloud-only and may not support on-premises deployment; they also have cold-start latency and execution time limits that conflict with low-latency inference requirements. Option D is wrong because deploying on a single powerful server creates a single point of failure and cannot scale horizontally to handle traffic spikes, leading to increased latency under load.

32
MCQmedium

A team is using an API from a cloud AI service to generate text. They notice that repeated requests with the same prompt return different outputs. They want consistent responses for testing. Which parameter should they adjust?

A.Increase the top_p parameter to 1.0
B.Set the frequency_penalty to 0
C.Increase the max_tokens parameter
D.Set the temperature to 0
AnswerD

Temperature controls sampling randomness; setting it to 0 makes the model select the highest-probability token at each step, producing near-deterministic output for identical prompts. This satisfies the requirement for consistent responses during testing, unlike top-p or penalty adjustments.

Why this answer

Temperature controls the randomness of the model's sampling distribution. Setting temperature to 0 makes the model deterministically pick the highest-probability token at each step, producing consistent (near-identical) outputs for the same prompt. This is the standard setting for reproducible testing.

Exam trap

AI0-001 often tests whether candidates confuse temperature (randomness) with top_p (nucleus sampling breadth) or max_tokens (length), leading them to pick top_p=1.0 as a determinism fix.

How to eliminate wrong answers

Option A is wrong because top_p=1.0 means consider all tokens in the nucleus, which actually increases diversity rather than reducing it. Option B is wrong because frequency_penalty=0 is the neutral default — it doesn't reduce randomness, it just stops penalizing repeated tokens. Option C is wrong because max_tokens only caps output length; it has no effect on determinism or sampling randomness.

33
MCQhard

A company is deploying a real-time object detection model on a fleet of IoT cameras. The model must run at 30 FPS on a device with limited memory and no internet connectivity. Which combination of techniques is MOST suitable?

A.Use FP16 inference and deploy via Docker containers
B.Use model distillation to create a smaller model and deploy via ONNX Runtime
C.Deploy on a GPU-based edge server with a full PyTorch model
D.Apply INT8 quantization and pruning, then deploy using TensorFlow Lite
AnswerD

INT8 quantization reduces memory footprint and accelerates inference; pruning removes redundant parameters. TensorFlow Lite is optimized for edge devices.

Why this answer

INT8 quantization reduces model size and latency, while pruning removes redundant weights, making the model suitable for memory-constrained edge devices. TensorFlow Lite is optimized for on-device inference with no internet dependency, supporting real-time 30 FPS object detection on IoT cameras.

Exam trap

CompTIA often tests the misconception that any lightweight deployment framework (like ONNX Runtime) is sufficient for edge devices, ignoring the need for hardware-specific quantization and pruning to meet strict memory and FPS constraints.

How to eliminate wrong answers

Option A is wrong because FP16 inference reduces precision but still requires significant memory and compute resources; Docker containers add overhead and are not designed for ultra-low-memory IoT cameras. Option B is wrong because model distillation creates a smaller model, but ONNX Runtime is a cross-platform inference engine that does not inherently provide the aggressive memory and latency optimizations needed for 30 FPS on constrained devices; it also lacks native support for hardware-specific quantization like TensorFlow Lite. Option C is wrong because deploying a full PyTorch model on a GPU-based edge server contradicts the 'limited memory and no internet connectivity' constraint; GPUs are power-hungry and expensive, and PyTorch's runtime overhead is too high for a memory-constrained IoT camera.

34
MCQmedium

An AI platform team is building a feature store that feeds both offline training jobs and an online model that must return features within a few milliseconds. They are concerned that a feature computed one way during training could be computed differently at serving time. Which design choice best prevents this training-serving skew?

A.Precompute all features nightly and have the online model read them from the same batch tables used for training.
B.Log raw request payloads during serving and retrain the model on that log so training data matches production inputs.
C.Define each feature once in a shared transformation definition that is executed by both the batch and online paths.
D.Compute features with separate code paths for batch training and online serving, and reconcile differences in periodic audits.
AnswerC

A single transformation definition executed by both the offline and online engines guarantees that the same logic produces training values and serving values, which removes the divergence that causes skew. The online engine materializes the result in a low-latency store, so millisecond serving requirements are met without duplicating the feature logic.

Why this answer

Training-serving skew arises when the same feature is computed by different logic in the offline and online environments. A shared transformation definition executed by both engines makes the computation identical by construction, and the online path simply materializes the result into a low-latency store, satisfying the millisecond requirement without duplicating logic.

Exam trap

The trap here is assuming that logging production data and retraining closes the gap, when skew comes from divergent feature logic rather than from differing input distributions.

35
MCQmedium

An MLOps engineer is building a feature pipeline for a recommendation model. Features must be served to the online model with single-digit millisecond latency, while the same feature definitions must also be usable by offline training jobs to prevent training-serving skew. Which component of a feature store architecture directly satisfies the low-latency serving requirement?

A.The feature registry that versions feature definitions
B.A scheduled batch job that materializes features every hour
C.The online store backed by a low-latency key-value database
D.The offline store backed by a data lakehouse
AnswerC

The online store is the serving layer of a feature store, typically implemented on a low-latency key-value database such as Redis or Bigtable. It holds the latest feature values keyed by entity ID and returns them in single-digit milliseconds during inference. Using the same feature definitions registered in the feature store for both online and offline paths is what prevents training-serving skew.

Why this answer

A feature store splits storage into an offline store for training and an online store for serving. The online store uses a low-latency key-value database so the model can retrieve the latest feature values in single-digit milliseconds. Because both paths are populated from the same registered feature definitions, the team also avoids training-serving skew, which is the second requirement in the scenario.

Exam trap

The trap here is assuming that a feature registry or a materialization job is the serving path, when only the online store actually answers low-latency inference reads.

36
MCQmedium

A media company runs an on-premises inference cluster for an image-tagging model. The model is trained on-premises and deployed into a container image that is rebuilt nightly in the company's internal registry. Security policy forbids any outbound internet access from the cluster. Which deployment approach best fits these constraints?

A.Export the trained model to ONNX, then serve it with ONNX Runtime inside the container image pulled from the internal registry.
B.Deploy the model to a managed cloud inference endpoint and route internal cluster traffic to it through the corporate VPN.
C.Store the model in a public Hugging Face repository and have the container download it at startup with a read-only token.
D.Use TensorFlow Serving with a remote model repository mounted over SMB from a partner network share.
AnswerA

ONNX Runtime is a self-contained inference engine that can be installed from internal package mirrors and loaded directly from the container image, so the model never has to fetch weights or metadata from the internet at runtime. Exporting to ONNX also fixes the computation graph ahead of time, which keeps nightly rebuilds deterministic and avoids any dependency on external model hubs.

Why this answer

An air-gapped cluster needs a self-contained inference stack: the runtime, the model, and all dependencies must travel inside the container image from the internal registry. Exporting to ONNX produces a portable, framework-independent graph that ONNX Runtime can execute without contacting any external service, so nightly rebuilds remain deterministic and no outbound traffic is required.

Exam trap

The trap here is assuming that a model must always be fetched from a hub or endpoint at runtime, when a self-contained ONNX artifact inside the image removes that dependency entirely.

37
MCQeasy

An organization wants to integrate an AI-powered summarization feature into their existing web application. The AI service will be called via API. Which factor is MOST important to consider for cost management?

A.Token pricing of the AI model
B.Authentication method (API key vs. OAuth)
C.Rate limits per minute
D.Network latency to the API endpoint
AnswerA

Token pricing directly governs API call costs: providers bill per input and output token, so summarisation of long documents scales expense with token volume. This satisfies the stem's API-based cost management constraint, since compute, storage and licensing are not consumed by the web application itself.

Why this answer

Token pricing directly determines the cost of each API call because AI models charge based on the number of input and output tokens processed. Since the summarization feature will make frequent API calls, token pricing is the primary cost driver. Other factors like authentication, rate limits, and latency affect security, throughput, and performance, but not the direct cost per request.

Exam trap

AI0-001 often tests the distinction between cost drivers and operational factors; candidates may incorrectly focus on rate limits or authentication as cost-related, but token pricing is the direct cost component.

How to eliminate wrong answers

Option B is wrong because authentication method affects security and implementation complexity, not the cost of using the AI service. Option C is wrong because rate limits control how many requests can be made per minute, which impacts scalability and throttling, but not the per-request cost. Option D is wrong because network latency affects response time and user experience, not the monetary cost of the AI service.

38
Multi-Selectmedium

A team is deploying a model that must comply with GDPR. Users can request deletion of their data. Which TWO practices should be implemented to support this compliance? (Select TWO.)

Select 2 answers
A.Enable output caching for frequently requested predictions
B.Validate inputs to prevent prompt injection attacks
C.Use a vector database to store user embeddings
D.Maintain the ability to delete a user's data from training sets and derived features
E.Implement data versioning and lineage tracking
AnswersD, E

Directly supports the right to erasure by allowing removal of user data and any features based on it.

Why this answer

GDPR's 'right to erasure' requires that upon user request, the organization must delete not only the user's raw data but also any derived features or embeddings that were generated from that data. Without this capability, the model could still indirectly retain user information through trained parameters or feature stores, violating compliance.

Exam trap

CompTIA often tests the misconception that simply using a vector database or caching mechanism satisfies GDPR deletion requirements, when in fact the critical practice is maintaining the ability to delete user data from all derived artifacts, including training sets and feature stores.

39
MCQmedium

A data scientist is using PyTorch to train a custom NLP model. The training is slow on a single GPU. They want to speed up training by using multiple GPUs on a single machine. Which PyTorch feature should they use?

A.TorchScript tracing
B.torch.nn.DataParallel
C.torch.optim.SGD
D.PyTorch Lightning's zero_grad function
AnswerB

DataParallel splits each batch across all GPUs on one machine and gathers gradients back to the primary device, giving single-process, multi-GPU acceleration with a one-line wrapper. It directly satisfies the single-machine, multiple-GPU constraint, unlike DistributedDataParallel, which targets multi-node scaling.

Why this answer

torch.nn.DataParallel is PyTorch's built-in module for single-machine, multi-GPU training. It wraps the model and automatically splits the input batch across available GPUs, replicating the model on each device and gathering outputs on the primary GPU — requiring only a one-line code change (wrapping the model).

Exam trap

AI0-001 often tests the confusion between DataParallel and DistributedDataParallel — candidates pick DDP for single-machine multi-GPU, but the question emphasizes minimal code changes, and DataParallel requires only wrapping the model, while DDP requires initializing a process group and launching multiple processes.

How to eliminate wrong answers

Option A (TorchScript tracing) is wrong because TorchScript is for serializing and optimizing models for deployment, not for multi-GPU training. Option C (torch.optim.SGD) is wrong because SGD is an optimizer, not a parallelism mechanism; it does not distribute work across GPUs. Option D (PyTorch Lightning's zero_grad function) is wrong because zero_grad is a gradient-clearing utility, not a multi-GPU training feature; PyTorch Lightning does support multi-GPU via its Trainer, but the zero_grad function itself is unrelated.

40
MCQmedium

A media company stores thousands of hours of raw broadcast footage in a cloud object storage bucket. A data engineering team needs a training dataset that contains only the short clips where a goal is scored, so they must locate and extract those specific time ranges from the video files before training. Which technology should the team use to extract the required segments from the video objects?

A.Use FFmpeg to decode the video and cut the required time ranges into new clip files.
B.Use a data lakehouse table format with schema evolution to store the video files as Delta tables.
C.Use OpenCV's VideoCapture with a GPU-accelerated codec to retrain the object detection model directly on the raw footage.
D.Use Apache Parquet to store each frame as a column and filter on the goal timestamp column.
AnswerA

FFmpeg is the standard open-source toolkit for demuxing, decoding, and re-encoding media streams, and it supports precise time-based trimming with parameters such as -ss and -to. Running it over the objects lets the team extract exactly the goal segments and write new clip files that become clean training data for the model pipeline.

Why this answer

Trimming video to precise time ranges is a media-processing task, and FFmpeg is the purpose-built tool for demuxing, decoding, and re-encoding streams with start and end timestamps. Object storage and table formats organize files but cannot parse video containers, so the extraction must be performed by a codec-aware utility before the clips enter the training pipeline.

Exam trap

The trap here is assuming that a data storage or table format can perform media manipulation, when extracting video segments requires a codec-aware media tool.

41
MCQmedium

A logistics company runs a route-optimization model on a fleet of delivery vehicles. Each vehicle has an NVIDIA Jetson module with limited memory, and connectivity is unavailable for hours at a time. The team wants the smallest possible runtime footprint while still executing the trained graph on the GPU. Which approach best fits these constraints?

A.Deploy the full training framework on each vehicle
B.Store the model as a CSV of weights and load it at startup
C.Serve the model from a cloud endpoint over the cellular network
D.Convert the model to a TensorRT engine for the target GPU
AnswerD

TensorRT builds a hardware-specific optimized engine from the trained graph, fusing layers, selecting tuned GPU kernels, and applying precision calibration. The resulting runtime footprint is far smaller than a full framework, and the engine is purpose-built for the Jetson GPU. Because the engine is generated ahead of time, vehicles can run it for hours without connectivity.

Why this answer

TensorRT compiles a trained network into a GPU-specific engine, fusing operations and choosing tuned kernels for the Jetson hardware. The engine is compact, runs entirely on the device, and needs no network, so the route optimizer keeps working during hours offline. Full frameworks, cloud endpoints, and raw weight files each fail one of the footprint, offline, or GPU-execution requirements.

Exam trap

The trap here is assuming any on-device copy of the model satisfies the constraint, when the requirement is a compact, GPU-optimized engine rather than a portable file or a cloud call.

42
MCQeasy

A hospital's radiology department wants to run a diagnostic imaging model inside its own data center because patient images cannot leave the premises. The team needs to manage model versions, roll back a bad deployment quickly, and keep an audit trail of which model version produced each prediction. Which approach best satisfies these requirements?

A.Copy the model file onto each radiologist workstation and run inference locally with a desktop script.
B.Host the model on a public inference API and send only anonymized image hashes for scoring.
C.Serve the model from an internal model registry with versioned artifacts and log the model version with each prediction.
D.Store the model in a shared network folder and have the serving application load whichever file has the newest timestamp.
AnswerC

An internal model registry keeps every artifact on-premises, assigns immutable versions, and lets the team promote or roll back a deployment by pointing the serving layer at a different version. Recording the version alongside each prediction creates the audit trail the department needs to trace any diagnostic output back to the exact model that produced it.

Why this answer

Keeping the model in an internal registry preserves data residency while giving the team immutable versions for promotion and rollback. Logging the model version with every prediction produces the traceability regulators and clinicians need, so a questionable result can be tied to a specific artifact rather than to an undifferentiated deployment.

Exam trap

The trap here is treating version control as simply keeping old files around, when the audit requirement demands that each prediction record the exact immutable version that produced it.

43
Multi-Selectmedium

A data science team wants to implement a feature store to serve pre-computed features for both training and inference with low latency. Which TWO tools are commonly used for building a feature store?

Select 2 answers
A.Kubeflow
B.Apache Hive
C.Feast
D.Tecton
E.MLflow
AnswersC, D

Feast is an open-source feature store that manages and serves features.

Why this answer

Feast (Feature Store) is an open-source operational data system that manages and serves machine learning features to both training and inference pipelines with low latency. It provides a consistent feature definition API, offline serving for training, and online serving via a low-latency store like Redis or DynamoDB, making it a standard choice for feature store implementations.

Exam trap

CompTIA often tests the distinction between ML orchestration tools (Kubeflow, MLflow) and dedicated feature stores (Feast, Tecton), trapping candidates who confuse lifecycle management with feature serving infrastructure.

44
Multi-Selecthard

A financial services company needs to deploy an ML model for loan approval that must be explainable to regulators. The model is a gradient boosting ensemble. They need to track experiments, log model parameters, and serve the model with explanations. Which THREE tools from the MLOps ecosystem should they use?

Select 3 answers
A.Apache Kafka
B.Docker Compose
C.Weights & Biases
D.Kubeflow
E.MLflow
AnswersC, D, E

W&B provides experiment logging, hyperparameter tracking, and model visualization.

Why this answer

Weights & Biases (W&B) is correct because it provides experiment tracking, hyperparameter logging, and model versioning, which are essential for regulatory explainability and auditability. It integrates directly with gradient boosting frameworks like XGBoost and LightGBM to log parameters and metrics, enabling reproducible ML pipelines. MLflow is correct because it offers experiment tracking, parameter and metric logging, a model registry, and model serving, allowing the team to track experiments and deploy the model with versioned artifacts.

Kubeflow is correct because it provides an end-to-end MLOps platform for building, training, and serving models on Kubernetes, including pipeline and model-serving components that support explainability tooling. Together, these three purpose-built MLOps tools cover experiment tracking, parameter logging, and model serving with explanations, whereas Apache Kafka (a streaming platform) and Docker Compose (a container orchestration tool for local development) do not address these MLOps requirements.

Exam trap

CompTIA AI+ often tests the distinction between general infrastructure tools (like Kafka or Docker Compose) and purpose-built MLOps tools (like W&B, Kubeflow, and MLflow) that directly address experiment tracking, model serving, and explainability.

45
MCQhard

A financial institution is designing an AI system to detect fraudulent transactions in real time. The system must process 10,000 transactions per second with sub-10 ms latency. The team plans to use a gradient boosting model. Which infrastructure component is most critical to meet the latency requirement?

A.A container orchestration platform like Kubernetes for auto-scaling.
B.A distributed streaming platform like Apache Kafka for data ingestion.
C.A GPU-accelerated database for storing transaction history.
D.A low-latency model serving framework that supports in-memory inference and batching.
AnswerD

To achieve sub-10 ms latency, the model serving framework must minimize overhead, support in-memory model execution, and efficiently batch requests. Frameworks like NVIDIA Triton Inference Server or TensorFlow Serving with optimized batching can deliver microsecond-level inference. This directly addresses the latency requirement by reducing per-request processing time.

Why this answer

Sub-10 ms latency requires a serving framework that minimizes overhead and efficiently handles concurrent requests. In-memory inference with optimized batching reduces per-request time, while other components like Kafka, databases, or Kubernetes add latency or do not directly affect inference speed. The serving layer is the bottleneck for meeting such strict latency.

Exam trap

The trap here is focusing on data ingestion or orchestration tools when the latency requirement is about the inference execution path.

46
MCQmedium

A data engineering team is building a pipeline to ingest streaming user activity data, process it in real-time, and store features in a feature store for ML models. Which streaming technology is BEST suited for this real-time data ingestion and processing?

A.Apache Kafka
B.Apache Spark SQL
C.Apache Airflow
D.Apache Hadoop MapReduce
AnswerA

Kafka provides high-throughput, fault-tolerant streaming for real-time data pipelines.

Why this answer

Apache Kafka is the best choice because it is a distributed streaming platform designed for high-throughput, fault-tolerant, real-time data ingestion and processing. It provides publish-subscribe messaging, durable log storage, and stream processing capabilities, making it ideal for ingesting streaming user activity data and feeding it into a feature store for ML models.

Exam trap

CompTIA often tests the distinction between batch and stream processing technologies, and the trap here is that candidates confuse Apache Spark SQL (a batch-oriented SQL engine) with Spark Streaming, or mistake Airflow's scheduling capabilities for real-time ingestion.

How to eliminate wrong answers

Option B (Apache Spark SQL) is wrong because Spark SQL is a module for structured data processing using SQL queries, not a streaming ingestion technology; while Spark Streaming exists, Spark SQL itself is not designed for real-time data ingestion. Option C (Apache Airflow) is wrong because Airflow is a workflow orchestration tool for batch scheduling and DAG management, not a real-time streaming ingestion or processing system. Option D (Apache Hadoop MapReduce) is wrong because MapReduce is a batch processing framework that processes data in large, static batches with high latency, making it unsuitable for real-time streaming ingestion.

47
Multi-Selectmedium

A healthcare analytics team is preparing to fine-tune a 7-billion-parameter open-weight language model on a single server with four NVIDIA A100 40 GB GPUs. Full fine-tuning runs out of memory, and the team wants to train on their clinical notes dataset while keeping GPU memory within the available budget. Which TWO techniques should they apply to reduce memory consumption during fine-tuning? (Choose two.)

Select 2 answers
A.Converting the model to ONNX format before training
B.Quantization-aware training of the base weights at 4-bit
C.Gradient checkpointing
D.LoRA (Low-Rank Adaptation)
E.Increasing the global batch size
AnswersC, D

Gradient checkpointing discards most intermediate activations during the forward pass and recomputes them during backpropagation. This trades additional compute for a large reduction in activation memory, which is a major consumer at long sequence lengths. Combined with parameter-efficient tuning, it helps fit the fine-tuning job into the four A100 40 GB GPUs available to this team.

Why this answer

Parameter-efficient tuning with LoRA removes the need to store optimizer state for all base parameters, and gradient checkpointing cuts activation memory by recomputing intermediates during the backward pass. Together they target the two dominant memory consumers in fine-tuning, making a 7B model trainable on four 40 GB GPUs without changing the base weights.

Exam trap

The trap here is assuming any low-precision or format-conversion step reduces training memory, when only the optimizer-state and activation reductions actually free GPU memory during fine-tuning.

48
Multi-Selectmedium

A data scientist needs to store large volumes of unstructured log data for future AI model training. They also need to run SQL-based analytics on the data. Which THREE services are appropriate for this requirement? (Choose 3)

Select 3 answers
A.Pinecone
B.Snowflake
C.BigQuery
D.Amazon S3
E.pgvector
AnswersB, C, D

Snowflake is a data warehouse that supports SQL analytics on structured/semi-structured data.

Why this answer

Snowflake is correct because it is a cloud-native data warehouse that supports both structured and semi-structured data (like JSON, Avro, Parquet) via its VARIANT data type, enabling SQL-based analytics on unstructured log data. It also integrates with cloud storage (e.g., Amazon S3) for storing large volumes of raw logs, making it suitable for AI model training pipelines.

Exam trap

CompTIA AI often tests the distinction between purpose-built databases (vector databases like Pinecone and pgvector) and general-purpose analytics platforms (Snowflake, BigQuery, S3), leading candidates to mistakenly select vector databases for log storage and SQL analytics.

49
MCQeasy

A hospital wants to run a natural language processing model that summarizes clinical notes. Because of patient privacy regulations, the data cannot leave the hospital's on-premises network, and there is no dedicated GPU available. Which deployment approach best fits these constraints?

A.Host the model in a colocation facility that is physically separate but connected by a dedicated VPN.
B.Fine-tune a large foundation model on the hospital's CPU servers and serve it with full FP32 precision.
C.Deploy a quantized open-source model on an on-premises CPU server using a runtime such as ONNX Runtime or llama.cpp.
D.Use a cloud-based LLM API and send de-identified notes for summarization.
AnswerC

Quantized open-source models can run entirely on-premises on CPU using runtimes like ONNX Runtime or llama.cpp, keeping patient data inside the network and requiring no GPU. Quantization reduces memory and compute demands enough for CPU inference on summarization tasks. This satisfies both the privacy constraint and the hardware limitation.

Why this answer

Running a quantized open-source model locally with a CPU-optimized runtime keeps all clinical data inside the hospital network and avoids GPU requirements. Quantization shrinks model size and speeds inference enough for summarization on standard servers. Cloud APIs, remote hosting, and full-precision large-model training all violate either the privacy boundary or the hardware constraint.

Exam trap

The trap here is treating de-identification as equivalent to keeping data on-premises, when the stated policy forbids any external transmission of clinical notes.

50
Multi-Selectmedium

A retail analytics team is building a retrieval-augmented generation assistant over product manuals. They need a vector index that supports fast approximate nearest neighbor search and can be updated as new manuals are published without rebuilding the entire index. Which TWO components should they use to meet these requirements? (Choose two.)

Select 2 answers
A.A columnar analytics warehouse that stores embeddings as arrays for SQL aggregation.
B.A message queue that streams manual PDFs directly into the index without transformation.
C.A dedicated vector database such as Milvus or Pinecone that supports incremental upserts and ANN indexes like HNSW.
D.A relational database with B-tree indexes on the manual text column.
E.An embedding model that converts each manual chunk into a dense vector before insertion.
AnswersC, E

Vector databases are built for embedding storage and similarity search, and they expose ANN index types such as HNSW that trade a small recall loss for large speed gains. Crucially, they support upserting individual vectors, so newly published manuals can be added without rebuilding the whole collection, which is exactly what the team needs.

Why this answer

A working retrieval pipeline needs embeddings to represent manual chunks semantically and a vector store that indexes those embeddings for fast approximate similarity search while allowing incremental additions. Combining an embedding model with a vector database that supports upserts and HNSW-style indexes satisfies both the speed and the update-without-rebuild requirements.

Exam trap

The trap here is assuming that any database capable of storing arrays can serve nearest neighbor queries efficiently, when only ANN-capable vector stores provide the required search speed.

51
MCQhard

A financial institution is deploying a real-time anomaly detection model on a Kubernetes cluster. The model must process streaming transactions with low latency and scale horizontally during peak hours. The team wants to use a serving solution that integrates natively with Kubernetes and supports autoscaling based on request concurrency. Which solution best meets these requirements?

A.KServe with Knative autoscaling
B.Standalone TensorFlow Serving deployed as a Kubernetes Deployment
C.NVIDIA Triton Inference Server with a fixed number of replicas
D.A custom Flask API wrapping the model, deployed with a Kubernetes Service
AnswerA

KServe integrates with Kubernetes and uses Knative to provide request-based autoscaling, including scale-to-zero and concurrency metrics. It supports low-latency inference and horizontal scaling, directly matching the streaming transaction requirements. This makes it the most appropriate choice.

Why this answer

KServe with Knative autoscaling provides native Kubernetes integration and request-concurrency-based scaling, which is essential for handling unpredictable transaction volumes with low latency. The other options either lack autoscaling, require manual configuration, or are not optimized for production inference.

Exam trap

The trap here is assuming that deploying any model server on Kubernetes automatically provides request-concurrency autoscaling, when it often requires additional components.

52
MCQmedium

A security team needs to ensure that all data used for AI model training in the cloud is encrypted at rest and in transit. Which set of measures meets this requirement on AWS?

A.Use Security Groups and Network ACLs
B.Use client-side encryption and store keys in AWS Secrets Manager
C.Enable S3 default encryption with SSE-S3 and use HTTPS for API calls
D.Enable VPC peering and use VPN connections
AnswerC

S3 default encryption with SSE-S3 applies AES-256 encryption to every object at rest automatically, satisfying the storage constraint without per-object configuration. HTTPS enforces TLS for data in transit during API calls. Together they cover both required states, though SSE-S3 lacks the customer-managed key control that SSE-KMS provides.

Why this answer

Encryption at rest is achieved by enabling default encryption on S3 (SSE-S3 or SSE-KMS), and encryption in transit is achieved by enforcing HTTPS/TLS for all API calls. Together, these two measures directly satisfy the requirement for data used in AI training on AWS.

Exam trap

AI0-001 often tests the distinction between network security controls (Security Groups, NACLs, VPN) and encryption controls; candidates mistakenly pick network options for encryption requirements.

How to eliminate wrong answers

Option A is wrong because Security Groups and NACLs control network traffic (firewall rules), not encryption of data at rest or in transit. Option B is wrong because client-side encryption with Secrets Manager only addresses encryption at rest for specific data and does not cover data in transit; it also adds complexity and is not a complete solution. Option D is wrong because VPC peering and VPN provide private network connectivity, not encryption of data at rest, and VPN encrypts transit but not storage.

53
MCQeasy

A data engineer is building a real-time feature store for an AI recommendation engine. The system must ingest millions of clickstream events per second, retain each event for 7 days, and allow the ML model to read the most recent user activity with sub-10ms latency. Which storage technology should the engineer select for the online feature serving layer?

A.Amazon S3 with Parquet files and AWS Glue catalog
B.Apache Kafka with a 7-day retention policy
C.A Redis in-memory key-value store with TTL-based expiration
D.A Snowflake data warehouse with a materialized view
AnswerC

Redis keeps feature values in memory and supports O(1) key lookups keyed by entity ID, which is exactly the access pattern an online feature store needs. Per-key TTL lets the engineer expire events after 7 days automatically, and the sub-millisecond in-memory read path comfortably satisfies the sub-10ms latency requirement even under high request concurrency.

Why this answer

Online feature serving requires a low-latency, key-value access pattern keyed by entity ID, with automatic expiry of stale features. An in-memory key-value store provides O(1) lookups in sub-millisecond time and supports TTL-based eviction, satisfying both the 7-day retention and sub-10ms latency requirements. Streaming platforms and analytical stores serve ingestion and batch computation roles instead.

Exam trap

The trap here is assuming that a streaming platform like Kafka, because it ingests real-time events, also serves them as a low-latency online feature store.

54
MCQmedium

A company wants to use a pre-trained model from a cloud-based AI service but must ensure that customer data is not used to improve the service. Which configuration should they choose?

A.Set data retention to 30 days
B.Enable content filtering
C.Use a data policy that prohibits training on customer data
D.Enable rate limiting
AnswerC

A data policy prohibiting training on customer data contractually and technically restricts the provider from using submitted inputs to improve or retrain the service, directly satisfying the requirement that customer data is not repurposed. This is the configuration control that enforces the no-training guarantee.

Why this answer

The 'No Training' data policy option explicitly prevents the cloud AI service from using customer prompts and completions to retrain or improve the underlying models. This configuration is essential for compliance with data privacy requirements, ensuring that customer data remains isolated from model improvement pipelines.

Exam trap

The trap here is that candidates often confuse data retention settings (which control storage duration) with data usage policies (which control whether data is used for training), leading them to select Option A instead of the correct data privacy policy that prevents training.

How to eliminate wrong answers

Option A is wrong because setting data retention to 30 days controls how long input and output data is stored for monitoring or debugging, but it does not prevent that data from being used for model training during that period. Option B is wrong because enabling content filtering only blocks harmful or policy-violating content from being generated; it has no effect on whether customer data is used to improve the service. Option D is wrong because rate limiting controls the number of API requests per time unit to manage load and cost, but it does not address data usage for training purposes.

55
MCQhard

An AI team is optimizing a convolutional neural network (CNN) for inference on a mobile device. The model has many layers and uses 32-bit floating-point weights. They need to reduce the model size and latency without significant accuracy loss. Which technique should they apply?

A.Pruning
B.Knowledge distillation
C.Quantization
D.Data augmentation
AnswerC

Quantization reduces the precision of the model's weights and activations, typically from 32-bit floating-point to 8-bit integers. This significantly decreases model size and speeds up inference on mobile hardware that supports integer operations. It can be done post-training or with quantization-aware training to minimize accuracy loss.

Why this answer

Quantization is the most effective technique to reduce model size and latency on mobile devices by converting 32-bit floating-point weights to lower precision, such as 8-bit integers. This leverages hardware acceleration for integer operations and reduces memory bandwidth. Pruning and knowledge distillation can also help but are not as directly targeted at precision reduction.

Exam trap

The trap here is assuming that any model compression technique will equally reduce latency, when in fact quantization specifically addresses precision and hardware acceleration.

56
Multi-Selectmedium

A retail analytics team is preparing a recommendation model for production. They need to serve many concurrent requests with predictable latency and also reduce the cost of running the model on GPU nodes. Which TWO practices best support these goals? (Choose two.)

Select 2 answers
A.Retrain the recommendation model with a larger embedding dimension to improve ranking quality.
B.Store the model on a network file share so all replicas read the same artifact at startup.
C.Increase the number of replicas so every request is handled by a dedicated pod with no queuing.
D.Batch multiple incoming requests into a single forward pass within a short time window.
E.Convert the model to a lower-precision format such as FP16 or INT8 before deployment.
AnswersD, E

Dynamic batching groups concurrent requests so the accelerator processes several inputs per forward pass, which raises throughput and spreads fixed per-pass overhead across more requests. On GPU nodes this directly lowers the cost per prediction, and because the batching window is bounded, latency stays predictable instead of growing with queue depth.

Why this answer

Throughput and cost on GPU nodes improve when each forward pass does more useful work and each operation is cheaper. Dynamic batching lets one pass serve several concurrent requests, while lower-precision execution reduces the cost of every operation. Together they raise requests served per GPU-second, which is the lever that lowers cost per prediction while keeping latency bounded.

Exam trap

The trap here is treating more replicas as the answer to concurrency, when adding accelerators increases cost instead of reducing it and leaves per-request efficiency unchanged.

57
MCQmedium

A team uses Apache Kafka to stream real-time sensor data for ML inference. They need to process the stream, perform feature engineering, and store results in a data lake. Which tool is best suited for this streaming ML pipeline?

A.Apache Spark with Structured Streaming
B.Apache Airflow
C.TensorFlow Data Validation
D.SageMaker Processing jobs
AnswerA

Structured Streaming runs the feature engineering as continuous incremental queries over Kafka topics, writing curated results to the data lake, and reuses Spark's ML libraries for inference. It handles both the streaming transformation and lake writes in one engine, unlike batch-only or queue-only alternatives.

Why this answer

Apache Spark with Structured Streaming is best suited because it provides a unified, scalable engine for both stream processing and batch processing, enabling real-time feature engineering on Kafka streams and direct writing to a data lake (e.g., Parquet format in Amazon S3). Its micro-batch or continuous processing model integrates natively with Kafka, allowing exactly-once semantics and low-latency transformations for ML inference pipelines.

Exam trap

CompTIA often tests the distinction between stream processing engines (like Spark Structured Streaming) and orchestration or batch tools (like Airflow or SageMaker Processing), trapping candidates who confuse workflow scheduling with real-time data processing.

How to eliminate wrong answers

Option B (Apache Airflow) is wrong because it is a workflow orchestration tool for scheduling and managing DAGs, not a stream processing engine; it cannot perform real-time feature engineering on Kafka streams. Option C (TensorFlow Data Validation) is wrong because it is designed for data validation and schema inference in static datasets or batch pipelines, not for continuous stream processing or feature engineering on live sensor data. Option D (SageMaker Processing jobs) is wrong because it is a batch processing service for data preprocessing and model evaluation on static datasets, lacking native support for streaming ingestion from Kafka or real-time feature computation.

58
MCQmedium

A media company wants to generate short video summaries from long recordings using a generative AI model. The model is hosted in the cloud, and the company needs to minimize cost while handling unpredictable traffic spikes. Which cloud service model is most appropriate?

A.Dedicated GPU virtual machine with a fixed hourly rate
B.Reserved instances with a one-year commitment
C.Serverless inference endpoint with automatic scaling
D.On-premises GPU cluster with a load balancer
AnswerC

Serverless inference endpoints scale automatically with traffic and charge only for the compute used during requests, which minimizes cost for unpredictable spikes. This model eliminates idle capacity costs and matches the variable demand of video summarization workloads.

Why this answer

A serverless inference endpoint with automatic scaling aligns cost with actual usage and handles unpredictable traffic without manual intervention. Dedicated VMs, on-premises clusters, and reserved instances all involve either idle costs or commitment risks that are suboptimal for variable demand.

Exam trap

The trap here is equating reserved instances with cost savings for all workloads, when they only benefit steady, predictable usage.

59
MCQhard

An ML platform team is running a recommendation model on a Kubernetes cluster with GPU nodes. During peak traffic, inference pods are frequently evicted and restarted, causing latency spikes. The team wants to reduce restart frequency and keep GPU utilization high without changing the model. Which combination of Kubernetes configuration changes should they apply?

A.Set resource requests equal to limits for GPU and memory, and assign a high-priority PriorityClass to the inference pods
B.Configure a PodDisruptionBudget with minAvailable set to zero and enable cluster autoscaler
C.Use a Vertical Pod Autoscaler in recommendation mode and set the pod restart policy to Always
D.Add a Horizontal Pod Autoscaler targeting CPU utilization and increase the pod replica count
AnswerA

Setting requests equal to limits gives the pods Guaranteed QoS, which makes the kubelet far less likely to evict them under node pressure, and it reserves the GPU and memory they need. A high-priority PriorityClass ensures that if the node does come under pressure, the scheduler and kubelet prefer evicting lower-priority workloads instead of the inference pods, directly reducing restart frequency and stabilizing latency.

Why this answer

Evictions under node pressure are mitigated by raising pod QoS and priority. Setting requests equal to limits yields Guaranteed QoS, which protects the pods' resource allocation, and a high PriorityClass makes the kubelet evict other workloads first. Autoscalers and disruption budgets influence capacity and voluntary disruptions but do not shield the pods from pressure-driven eviction on a contended GPU node.

Exam trap

The trap here is assuming that adding more replicas or an autoscaler prevents evictions, when eviction is governed by pod QoS class and priority rather than replica count.

60
MCQeasy

A hospital wants to run a diagnostic image classifier entirely inside its own data center because patient images cannot leave the premises. The IT team needs a deployment model that keeps all data and inference local while still allowing the AI team to push updated model versions. Which deployment approach fits these requirements?

A.Deploy the model on on-premises GPU servers and manage updates with a self-hosted model registry and serving stack.
B.Use a public cloud inference endpoint with a private VPC connection from the hospital.
C.Run the classifier on each radiologist's workstation using a CPU-only runtime.
D.Use a serverless function in the cloud that processes images uploaded through a signed URL.
AnswerA

Running the classifier on local GPU hardware keeps every image and prediction inside the hospital network, satisfying the data residency constraint. A self-hosted registry and serving layer let the AI team version models and roll out updates internally, so the organization retains both control of the data and a workable release process.

Why this answer

Data residency demands that inference run where the images already reside, so on-premises GPU servers hosting the model are required. Pairing them with a self-hosted registry and serving stack preserves the ability to version and update models without sending patient data to any external environment.

Exam trap

The trap here is treating a private network path or encrypted upload as equivalent to keeping data on premises, when the decisive factor is where inference actually executes.

61
MCQmedium

A company is deploying a computer vision model to smartphones for offline object detection. The model was trained in PyTorch. Which format should they use for deployment on iOS devices?

A.TorchScript
B.ONNX
C.Core ML
D.TensorFlow Lite
AnswerC

Core ML is Apple's on-device inference framework, so converting the trained PyTorch model to Core ML format lets it run natively and offline on iOS hardware, using Neural Engine acceleration without a network connection or server round trip.

Why this answer

Core ML is Apple's native machine learning framework for iOS, macOS, and other Apple platforms, and it requires models in the Core ML format (.mlmodel). Converting a PyTorch model to Core ML (via coremltools) enables on-device inference with optimized performance and integration with Apple's hardware accelerators.

Exam trap

AI0-001 often tests the confusion between cross-platform formats (ONNX, TorchScript) and platform-native formats (Core ML for iOS, TensorFlow Lite for Android) — candidates who pick ONNX for portability miss that iOS requires Core ML.

How to eliminate wrong answers

Option A is wrong because TorchScript is a PyTorch serialization format for deploying models in PyTorch runtimes (e.g., LibTorch), not for iOS deployment. Option B is wrong because ONNX is an open interchange format supported by many runtimes, but iOS does not natively run ONNX models — it requires conversion to Core ML. Option D is wrong because TensorFlow Lite is Google's format for Android and edge devices, not the native format for iOS.

62
MCQhard

A media company uses a large language model (LLM) to generate article summaries. They want to reduce inference costs and latency without significantly degrading summary quality. The LLM is currently served at full precision. Which optimization technique is most appropriate?

A.Cache all generated summaries to avoid repeated inference.
B.Apply quantization to convert the model weights to lower precision (e.g., INT8).
C.Use a larger model with more parameters to improve summary quality.
D.Increase the batch size for inference requests.
AnswerB

Quantization reduces the precision of model weights and activations, typically from FP32 to INT8, which decreases memory usage and speeds up inference on compatible hardware. For LLMs, post-training quantization can significantly lower latency and cost with minimal quality loss, especially when using techniques like GPTQ or AWQ. This directly addresses the media company's need to optimize inference without major degradation.

Why this answer

Quantization converts model weights to lower precision, reducing memory bandwidth and compute requirements, which lowers latency and cost. For LLMs, INT8 quantization often preserves summary quality well. Increasing batch size helps throughput but not per-request latency, a larger model worsens cost, and caching is ineffective for unique articles.

Thus, quantization is the most appropriate optimization.

Exam trap

The trap here is confusing throughput optimizations like batching with latency and cost reductions, or assuming that caching will solve the problem when inputs are largely unique.

63
Multi-Selecteasy

A data scientist wants to develop a computer vision model using transfer learning. They need a framework that provides pre-trained models and easy-to-use APIs for data augmentation and training. Which TWO frameworks are best suited for this task?

Select 2 answers
A.Hugging Face Transformers
B.PyTorch
C.scikit-learn
D.TensorFlow
E.Keras
AnswersB, D

PyTorch provides torchvision with pre-trained models and torchvision.transforms for data augmentation, making it ideal for transfer learning in computer vision.

Why this answer

PyTorch (option B) is correct because it offers a rich ecosystem of pre-trained models via `torchvision.models`, along with built-in data augmentation transforms in `torchvision.transforms` and a flexible training loop that is ideal for transfer learning. Its dynamic computation graph makes it easy to modify model architectures for fine-tuning, which is a core requirement for the task.

Exam trap

Candidates often select Hugging Face Transformers because it provides pre-trained models, but it is primarily designed for NLP tasks, not computer vision. Similarly, Keras is a high-level API that runs on top of TensorFlow, so it is not considered a standalone framework for this purpose.

64
MCQmedium

A hospital's AI team is deploying a real-time patient deterioration prediction model on bedside monitoring devices. The devices have limited RAM (512 MB) and no GPU, and the model must perform inference within 50 ms. The team has a trained TensorFlow model saved as a SavedModel. Which deployment approach best meets these constraints?

A.Convert the SavedModel to TensorFlow Lite and apply post-training quantization to INT8.
B.Deploy the SavedModel using TensorFlow Serving on a central GPU server and stream patient data to it.
C.Use the SavedModel directly in a Python script with TensorFlow's default runtime on the bedside device.
D.Convert the model to ONNX and run it with ONNX Runtime on the bedside device without quantization.
AnswerA

TensorFlow Lite is designed for resource-constrained edge devices, and post-training INT8 quantization reduces model size and memory usage while accelerating inference on CPUs. This directly addresses the 512 MB RAM limit and 50 ms latency requirement without needing a GPU, making it the most suitable approach for bedside monitors.

Why this answer

TensorFlow Lite is optimized for edge devices with limited compute and memory. Post-training INT8 quantization reduces the model size by up to 4x and speeds up CPU inference, directly addressing the 512 MB RAM and 50 ms latency constraints. Other options either rely on external servers or fail to optimize the model for the device's limitations.

Exam trap

The trap here is assuming that any optimized runtime like ONNX Runtime automatically solves memory and latency issues without quantization.

65
MCQeasy

A developer wants to integrate an AI-powered text summarization API into their application. They need to authenticate securely and manage usage limits. What is the standard mechanism for authenticating with cloud-based AI services?

A.Provide a username and password in the request body
B.Embed the API key in the URL query string
C.Use a digital certificate for each request
D.Include an API key in the HTTP request header
AnswerD

API keys passed in the HTTP request header authenticate the calling application to the cloud AI service and let the provider meter and enforce usage quotas per key, directly satisfying the secure authentication and usage-limit management requirements in the stem.

Why this answer

Cloud-based AI services, including text summarization APIs, standardize authentication via API keys passed in the HTTP header (e.g., `Authorization: Bearer <key>` or `x-api-key: <key>`). This method keeps credentials out of URLs and request bodies, preventing exposure in logs or caches, and aligns with RESTful API best practices and OWASP guidelines for secure API access.

Exam trap

CompTIA often tests the misconception that embedding credentials in a URL or request body is acceptable for simplicity, but the trap here is that API keys must never appear in URLs or bodies due to security risks like exposure in server logs and referrer headers, making the HTTP header the only standard and secure option.

How to eliminate wrong answers

Option A is wrong because sending a username and password in the request body violates security best practices—credentials would be exposed in plaintext in logs, monitoring tools, and intermediate proxies, and it does not support stateless, token-based authentication used by modern AI APIs. Option B is wrong because embedding an API key in the URL query string exposes the key in server logs, browser history, and referrer headers, making it vulnerable to interception and violating RFC 3986 recommendations against sensitive data in URIs. Option C is wrong because digital certificates (e.g., mTLS) are typically used for machine-to-machine authentication in high-security enterprise environments, not as the standard mechanism for cloud AI services, which rely on simpler API key or OAuth 2.0 token flows for scalability and ease of integration.

66
MCQeasy

A data engineer needs to process streaming clickstream data for real-time feature engineering in an ML pipeline. Which data pipeline technology is BEST suited for this task?

A.Apache Spark in batch mode
B.Snowflake
C.Apache Kafka
D.Apache Airflow
AnswerC

Kafka is a distributed publish-subscribe log that ingests continuous event streams with low latency and durable ordering, satisfying the real-time feature engineering requirement. Batch stores such as S3 or HDFS cannot process clickstream events as they arrive, so they fail the streaming constraint.

Why this answer

Apache Kafka is the best choice because it is a distributed streaming platform designed for high-throughput, fault-tolerant, real-time data ingestion and processing. It can capture clickstream events as they occur and make them immediately available for feature engineering in an ML pipeline, supporting exactly-once semantics and low-latency delivery.

Exam trap

CompTIA AI often tests the distinction between data ingestion/messaging systems (Kafka) and batch processing or storage systems, leading candidates to confuse Airflow's orchestration role with actual stream processing capabilities.

How to eliminate wrong answers

Option A is wrong because Apache Spark in batch mode processes data in static, finite batches with high latency, making it unsuitable for real-time streaming clickstream data. Option B is wrong because Snowflake is a cloud-based data warehouse optimized for analytical queries on structured, stored data, not for real-time stream ingestion or processing. Option D is wrong because Apache Airflow is a workflow orchestration tool for scheduling and monitoring batch jobs, not a stream processing or messaging system capable of handling real-time data streams.

67
MCQhard

A media company trains a video tagging model on a large dataset in the cloud. The model will run inference on-premises in a facility with intermittent network connectivity, and the operations team wants to avoid re-authoring the model for each target runtime. Which deployment artifact best meets these constraints?

A.A pickle file of the scikit-learn estimator
B.An ONNX model executed with the ONNX Runtime
C.The native TensorFlow SavedModel directory
D.A Docker image containing the full training environment
AnswerB

ONNX is an open, framework-neutral model representation, and ONNX Runtime provides a portable inference engine that runs on Windows, Linux, and edge devices without the original training framework. Exporting the trained model to ONNX lets the same artifact execute on-premises despite intermittent connectivity, and the team avoids re-authoring the model for each target runtime because the graph format is standardized.

Why this answer

ONNX defines a standardized computation graph, and ONNX Runtime executes that graph across platforms and hardware without requiring the original training framework. Exporting to ONNX gives the media company a single artifact that runs on-premises in a disconnected facility and removes the need to re-author the model for each inference runtime, satisfying both stated constraints.

Exam trap

The trap here is equating containerization with runtime portability, when a container still embeds one specific framework build and does not standardize the model graph.

68
MCQeasy

A startup is building a conversational AI assistant that must understand and generate human-like text. The team has limited labeled data and a modest budget for compute. They want to leverage existing large language models rather than pretraining one. Which approach best meets their needs?

A.Use a traditional n-gram language model trained on the startup's text data to generate responses.
B.Use a pretrained large language model via an API or open-source checkpoint and apply prompt engineering or lightweight fine-tuning for the assistant's domain.
C.Train a transformer model from scratch on the startup's proprietary conversation logs to ensure full control over the architecture.
D.Implement a rule-based chatbot using regular expressions and decision trees to handle user intents.
AnswerB

Leveraging a pretrained LLM avoids the enormous cost of pretraining. Prompt engineering requires no parameter updates, and lightweight fine-tuning methods like LoRA or adapter tuning adapt the model to the domain with minimal compute and data. This matches the startup's constraints while still delivering human-like understanding and generation, making it the most practical and cost-effective path.

Why this answer

For a startup with limited data and compute, the most effective strategy is to build on a pretrained large language model rather than starting from scratch. Pretrained LLMs already possess broad language understanding and generation capabilities. Prompt engineering can steer behavior without training, and parameter-efficient fine-tuning methods like LoRA adapt the model to the domain with minimal resources.

This balances quality, cost, and speed to deployment.

Exam trap

The trap here is assuming that training from scratch or using classical NLP methods can match the language understanding of pretrained LLMs, when the startup's constraints make leveraging pretrained models the only viable path.

69
MCQmedium

A company has a TensorFlow model trained on-premises and wants to deploy it on AWS SageMaker for scalable inference. What is the BEST way to package the model for deployment?

A.Convert the model to ONNX and upload to SageMaker
B.Upload the .h5 file to S3 and create a SageMaker endpoint directly
C.Package the model in a Docker container with a TensorFlow serving script and push to Amazon ECR
D.Use SageMaker Studio to train the model again from scratch
AnswerC

SageMaker deploys models from container images in Amazon ECR, so packaging the TensorFlow model with a serving script inside a Docker image provides the inference stack SageMaker requires. Plain model artefacts or notebooks cannot be served directly.

Why this answer

SageMaker expects models in a container format; the inference container should include the model artifacts and the serving code, allowing SageMaker to host it on scalable endpoints.

70
MCQhard

During inference, a model served via a REST API occasionally returns high latency due to cold starts. The team uses a containerized service on Kubernetes with horizontal pod autoscaling. Which solution minimizes cold start impact while controlling cost?

A.Configure the autoscaler based on request count with a shorter cooldown period
B.Increase CPU and memory requests for the inference container
C.Switch to vertical pod autoscaling
D.Use a sidecar container that pre-warms the model and set a minimum replica count
AnswerD

Pre-warming ensures the model is loaded; minimum replicas keep pods ready, reducing cold starts.

Why this answer

A sidecar warm-up agent and a minimum replica count keep pods ready. Increasing resources may not fix cold starts; autoscaling based on request count may lag; vertical scaling helps but not directly.

71
Multi-Selectmedium

A data engineering team is designing a data pipeline to process streaming sensor data and feed it into an ML model for anomaly detection. Which THREE components are essential for this pipeline?

Select 3 answers
A.Apache Airflow for scheduling recurring batch jobs
B.Amazon S3 as a data lake for storing raw sensor data
C.Snowflake as a real-time streaming destination
D.Apache Kafka for ingesting streaming sensor data
E.Apache Spark Structured Streaming for real-time processing
AnswersB, D, E

S3 is a scalable object store that can serve as a data lake for raw sensor data, accessible for both streaming and batch processing.

Why this answer

Amazon S3 is essential as a data lake for storing raw sensor data because it provides durable, scalable, and cost-effective object storage that can serve as a central repository for streaming data before and after processing. In a streaming pipeline, raw data must be persisted for reprocessing, historical analysis, and compliance, and S3's integration with Apache Spark and Kafka makes it a natural landing zone for sensor data.

Exam trap

CompTIA often tests the distinction between batch and streaming technologies, and the trap here is that candidates confuse Airflow's scheduling capability with real-time streaming orchestration, or assume Snowflake can act as a streaming sink when it is fundamentally a batch-oriented warehouse.

72
Multi-Selectmedium

A machine learning engineer wants to track hyperparameter experiments and compare results across runs. Which TWO tools are best suited for this purpose? (Choose 2)

Select 2 answers
A.MLflow
B.Weights & Biases
C.Apache Airflow
D.Docker
E.Kubeflow
AnswersA, B

MLflow provides experiment tracking that logs parameters, metrics and artefacts per run, letting engineers compare hyperparameter configurations side by side. Its tracking server and UI directly satisfy the requirement to record and contrast results across multiple runs.

Why this answer

MLflow is correct because it provides a centralized tracking server and API to log hyperparameters, metrics, and artifacts for each run, enabling easy comparison across experiments. Weights & Biases is correct because it offers a cloud-hosted dashboard with real-time logging, hyperparameter sweeps, and collaborative comparison features, making it ideal for tracking and comparing runs.

Exam trap

CompTIA AI exams often test the distinction between infrastructure tools (orchestration, containerization) and purpose-built experiment tracking tools; the trap here is that candidates may confuse Kubeflow’s pipeline capabilities with dedicated experiment tracking, or assume Docker/Airflow can serve as tracking solutions because they are used in ML workflows.

73
MCQhard

A financial institution runs a credit-scoring model that must comply with internal governance requiring that every individual prediction be traceable to the input features that drove it, and that the explanation be produced at inference time for each applicant. The model is a complex gradient-boosted ensemble. Which approach best satisfies the requirement to generate a per-prediction explanation for each applicant?

A.Report feature importances from the trained ensemble
B.Apply SHAP values to each individual prediction
C.Increase the number of boosting rounds to improve model stability
D.Train a global surrogate decision tree on the ensemble's outputs
AnswerB

SHAP assigns each feature a contribution to a single prediction based on cooperative game theory, so every applicant receives an explanation showing which input values pushed the score up or down. It works with tree ensembles through efficient exact algorithms, and its additive guarantees make the per-prediction attribution auditable, matching the governance requirement precisely.

Why this answer

Per-prediction traceability requires a local explanation method that decomposes an individual score into feature contributions. SHAP provides exactly that by attributing the difference between a prediction and a baseline to each input feature, with efficient exact computation available for tree ensembles. Global summaries such as feature importance or surrogate trees describe aggregate behavior and cannot explain a single applicant's decision.

Exam trap

The trap here is treating global interpretability artifacts, such as feature importance or a surrogate tree, as if they explain individual predictions.

74
MCQhard

A research lab is training a large language model on a cluster of GPUs. They notice that training throughput decreases significantly when scaling from 8 to 16 GPUs. The model uses data parallelism with synchronous updates. Which factor is most likely causing the decreased throughput?

A.The learning rate is too high for the larger effective batch size.
B.The model is not using mixed precision training.
C.Insufficient GPU memory causing out-of-memory errors.
D.Increased communication overhead for gradient synchronization across GPUs.
AnswerD

In synchronous data parallelism, gradients must be averaged across all GPUs after each backward pass. As the number of GPUs increases, the all-reduce communication cost grows, potentially becoming a bottleneck. This overhead can reduce throughput if the network bandwidth or latency is insufficient, especially when scaling from 8 to 16 GPUs.

Why this answer

Synchronous data parallelism requires all-reduce operations to synchronize gradients. As the number of GPUs increases, the communication volume and frequency grow, and if the interconnect bandwidth is limited, this becomes a bottleneck. This is a common scaling challenge in distributed training, leading to sublinear speedup or even decreased throughput.

Exam trap

The trap here is attributing throughput drops to model or hyperparameter issues rather than inter-GPU communication overhead.

75
MCQeasy

A developer wants to deploy a scikit-learn model as a REST API endpoint with minimal infrastructure management. Which cloud service is MOST appropriate?

A.Use AWS Lambda with a custom runtime
B.Deploy on an EC2 instance manually
C.Use AWS SageMaker to create a real-time endpoint
D.Use Amazon ECS with manual Docker setup
AnswerC

SageMaker offers managed inference endpoints with automatic scaling, reducing operational overhead.

Why this answer

AWS SageMaker provides a fully managed service for deploying machine learning models as real-time endpoints with built-in scaling, monitoring, and automatic infrastructure management. It directly supports scikit-learn models via pre-built containers, eliminating the need for custom runtime setup or manual server configuration. This makes it the most appropriate choice for a developer seeking minimal infrastructure management.

Exam trap

CompTIA often tests the misconception that serverless compute like AWS Lambda is the best choice for any API deployment, but the trap here is that Lambda's execution environment and constraints (timeout, payload size, cold starts) make it inappropriate for ML model inference, whereas SageMaker is purpose-built for this workload.

How to eliminate wrong answers

Option A is wrong because AWS Lambda with a custom runtime requires manual packaging of the scikit-learn model and dependencies, and Lambda has a 15-minute timeout and limited memory, making it unsuitable for real-time inference with larger models or payloads. Option B is wrong because deploying on an EC2 instance manually involves provisioning, patching, scaling, and managing the underlying server, which contradicts the requirement for minimal infrastructure management. Option D is wrong because Amazon ECS with manual Docker setup still requires managing the cluster, task definitions, and scaling policies, adding operational overhead compared to SageMaker's fully managed endpoint service.

Page 1 of 2 · 114 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Aio Ai Infrastructure questions.