Courseiva

CCNA AI Infrastructure and Technologies Questions

39 of 114 questions · Page 2/2 · AI Infrastructure and Technologies · Answers revealed

76
Multi-Selecthard

A data science team uses Vertex AI for model training and deployment. They want to implement CI/CD for ML pipelines. Which THREE Google Cloud services should they integrate?

Select 3 answers
A.Vertex AI Pipelines
B.Cloud Deploy
C.BigQuery
D.Cloud Build
E.Google Kubernetes Engine (GKE)
AnswersA, B, D

Vertex AI Pipelines orchestrates the ML workflow itself, running training, evaluation and deployment steps as a reproducible DAG. It supplies the pipeline automation the CI/CD requirement demands, letting each code commit trigger retraining and validation before Cloud Deploy handles release promotion.

Why this answer

Vertex AI Pipelines (A) is correct because it orchestrates and automates the ML workflow steps (data prep, training, evaluation, deployment) as reproducible pipeline runs, which is the core of CI/CD for ML. Cloud Deploy (B) is correct because it provides managed continuous delivery to targets such as GKE, Cloud Run, and Anthos, enabling progressive rollout and approval gates for the deployment stage of the ML pipeline. Cloud Build (D) is correct because it executes the CI portion—building container images, running tests, and triggering pipeline jobs—and integrates natively with Vertex AI and Cloud Deploy via triggers and build steps.

BigQuery (C) is not correct here because it is a data warehouse/analytics service, not a CI/CD component, even though it may store training data. Google Kubernetes Engine (E) is not correct because it is a runtime platform for containers, not a CI/CD service, and Cloud Deploy can target it without GKE itself being the CI/CD integration.

Exam trap

The trap is selecting data or infrastructure services like BigQuery or GKE instead of the specific CI/CD services; candidates must recognize that CI/CD for ML requires build, orchestrate, and deploy tools, not data warehouses or container platforms.

77
MCQmedium

An AI platform team is deploying a large language model for internal document summarization. Legal requires that no prompt or document content leaves the company's virtual private cloud, and the security team wants to control the exact model weights and runtime version. The team already has GPU capacity reserved in their own VPC. Which deployment approach best satisfies these constraints?

A.Use a provider's managed private endpoint with a business associate agreement
B.Self-host the model weights on the reserved GPUs inside the VPC and expose an internal inference endpoint
C.Deploy the model to a serverless inference service in a different cloud region
D.Call a public foundation model API using customer-managed encryption keys
AnswerB

Self-hosting the weights on reserved GPUs keeps all prompt and document data inside the company VPC, satisfying the legal data-residency constraint. The team also controls the exact weight revision and runtime version, which meets the security requirement. Because GPU capacity is already reserved, the incremental cost is operational rather than a new capital outlay, making this approach both compliant and practical.

Why this answer

The constraints combine data residency, model weight control, and runtime version control, and the team already owns reserved GPU capacity. Running the weights on that capacity inside the VPC keeps all payloads private and lets the team pin both the weight revision and the serving runtime. Managed APIs and serverless services, even with private endpoints or encryption, still process content on provider infrastructure.

Exam trap

The trap here is treating a private network endpoint or customer-managed encryption key as equivalent to keeping data and model execution inside your own VPC.

78
MCQmedium

A team uses Kubeflow to manage ML workflows on Kubernetes. They want to automate hyperparameter tuning for a training job. Which Kubeflow component should they use?

A.KFServing
B.Kubeflow Notebooks
C.Kubeflow Pipelines
D.Kubeflow Katib
AnswerD

Katib is Kubeflow's dedicated hyperparameter tuning and neural architecture search component, running trials as Kubernetes jobs and applying algorithms such as Bayesian optimisation or random search. It satisfies the stem's automation requirement by launching and comparing training runs without manual intervention.

Why this answer

Kubeflow Katib is the dedicated component for automated hyperparameter tuning and neural architecture search in Kubeflow. It supports various search algorithms (e.g., Bayesian optimization, random search) and early stopping, and integrates natively with Kubernetes to run trials as parallel jobs. The team can define a hyperparameter search space and objective metric, and Katib will orchestrate the tuning process.

This directly addresses the requirement to automate hyperparameter tuning for a training job.

Exam trap

AI0-001 often tests the confusion between orchestration (Kubeflow Pipelines) and specialized tuning (Katib), causing candidates to pick Pipelines when asked about hyperparameter tuning.

How to eliminate wrong answers

Option A is wrong because KFServing (now KServe) is a model serving component for deploying and scaling inference services, not for hyperparameter tuning. Option B is wrong because Kubeflow Notebooks provides interactive Jupyter notebook environments for development and experimentation, but does not automate hyperparameter tuning. Option C is wrong because Kubeflow Pipelines is a workflow orchestration tool for defining and running ML pipelines, but it does not include built-in hyperparameter tuning algorithms; while it can be used to orchestrate Katib, it is not the component that performs the tuning itself.

79
MCQmedium

A hospital's radiology department is deploying an AI system that analyzes chest X-rays to flag potential pneumonia. Because patient data cannot leave the hospital's on-premises network, the model must run locally. The IT team wants to ensure the model's inference results can be explained to radiologists and auditors. Which approach best satisfies the explainability requirement while keeping the model on-premises?

A.Use a black-box deep learning model and provide a confidence score for each prediction.
B.Apply LIME or SHAP to generate local explanations for each prediction without modifying the model.
C.Deploy the model to a cloud service that offers built-in explainability dashboards.
D.Replace the deep learning model with a logistic regression classifier trained on the same data.
AnswerB

LIME and SHAP are post-hoc explainability techniques that approximate how a model's features contribute to individual predictions. They work with any black-box model and can run entirely on-premises, satisfying the data residency constraint. Radiologists can see which pixels or regions influenced a flag, and auditors gain documentation for compliance. This directly meets the requirement without retraining or altering the deployed model.

Why this answer

Post-hoc explainability techniques such as LIME and SHAP allow clinicians to understand individual predictions from complex models without sacrificing accuracy. They operate locally and do not require moving data off-premises, satisfying both the explainability and data residency requirements. Retraining a simpler model risks accuracy, while confidence scores and cloud dashboards fail to provide the needed transparency under the given constraints.

Exam trap

The trap here is assuming that a confidence score or a simpler model is sufficient for explainability, when the scenario demands insight into feature contributions while preserving model performance and on-premises data handling.

80
MCQmedium

A company is building a recommendation system that uses user embeddings stored in a vector database. The system must retrieve the top 10 most similar items for a given user query. Which vector database feature is MOST critical for this task?

A.Built-in data versioning
B.ACID transaction support
C.Approximate nearest neighbor (ANN) search
D.SQL query interface
AnswerC

Approximate nearest neighbour search indexes embeddings so the top 10 most similar items are retrieved without comparing every vector. This satisfies the low-latency similarity requirement, which exact brute-force comparison across a large embedding store cannot meet at scale.

Why this answer

Approximate nearest neighbor (ANN) search is the most critical feature because it enables the vector database to efficiently find the top-10 most similar items to a user query embedding without scanning the entire dataset. Unlike exact nearest neighbor search, ANN algorithms (e.g., HNSW, IVF) trade a small amount of accuracy for massive performance gains, which is essential for real-time recommendation systems handling millions of high-dimensional vectors.

Exam trap

CompTIA often tests the misconception that SQL or ACID features are needed for all database tasks, but in vector databases, the critical differentiator is the ANN search algorithm, not traditional relational or transactional capabilities.

How to eliminate wrong answers

Option A is wrong because built-in data versioning manages historical changes to data but does not directly impact the speed or accuracy of similarity search; it is irrelevant to the core retrieval task. Option B is wrong because ACID transaction support ensures data consistency and reliability during writes but does not optimize or accelerate vector similarity queries; it addresses transactional integrity, not search performance. Option D is wrong because a SQL query interface is designed for structured relational queries and lacks native support for high-dimensional vector similarity operations; using SQL for nearest neighbor search would require inefficient full-table scans or custom extensions, defeating the purpose of a vector database.

81
MCQmedium

A team is using a cloud AI service with a pay-per-token pricing model. They want to minimize costs while maintaining response quality. Which strategy is MOST effective?

A.Switch to a smaller, less capable model
B.Increase the batch size for API calls
C.Use prompt caching for repeated query patterns
D.Reduce the model's max_tokens to a very low value
AnswerC

Prompt caching stores previously processed prompt prefixes so repeated query patterns reuse cached context rather than resending and reprocessing full tokens. Since billing is per token, this directly reduces the tokens charged while preserving identical response quality, satisfying the stem's cost-minimisation constraint.

Why this answer

Prompt caching reduces costs by avoiding redundant token processing for repeated query patterns. The cloud AI service charges per token, so caching the prefix of frequent requests (e.g., system prompts or common context) means only the new, unique tokens are billed, directly lowering expenditure without sacrificing response quality.

Exam trap

Candidates often mistakenly think that reducing model size or output length is the only way to cut costs, but the correct strategy leverages architectural features like prompt caching to reduce token consumption without affecting quality.

How to eliminate wrong answers

Option A is wrong because switching to a smaller, less capable model typically reduces response quality, which contradicts the requirement to maintain quality. Option B is wrong because increasing batch size for API calls does not reduce per-token cost; it may improve throughput but still charges for all tokens processed. Option D is wrong because reducing max_tokens to a very low value can truncate responses, degrading quality, and does not address the cost of input tokens or repeated patterns.

82
MCQhard

An AI platform team runs inference for an image classifier on a shared GPU node. Multiple model replicas currently load the full model weights into GPU memory independently, and the node runs out of GPU memory when a third replica starts. The team wants to serve more replicas per GPU without changing model accuracy. Which approach best addresses the constraint?

A.Pin each replica to a separate CUDA stream and rely on the driver's scheduler to time-slice GPU memory.
B.Serve the replicas from shared GPU memory using a model server that supports weight sharing or a runtime that loads the model once.
C.Convert the model to FP16 and run the replicas with a smaller batch size to fit more instances.
D.Enable CUDA Unified Memory so the driver pages weights between host RAM and GPU memory on demand.
AnswerB

Frameworks such as NVIDIA Triton Inference Server and runtimes like NVIDIA TensorRT-LLM or vLLM allow multiple model instances to share a single copy of the weights in GPU memory, so additional replicas consume only activation and KV-cache space. This raises replica density on the same GPU while leaving the model's numerical precision and accuracy untouched, exactly matching the requirement.

Why this answer

The exhaustion comes from duplicate copies of identical weights, so the fix is to load the weights once and let multiple replicas reference them. Model servers and optimized runtimes that support shared weights or a single loaded model instance cut per-replica memory to activations and cache, increasing density without touching precision or accuracy.

Exam trap

The trap here is reaching for a precision change or memory paging to save space, when the actual waste is duplicated weight copies that should be shared.

83
MCQmedium

A company is using Google Cloud Vertex AI for model training. They want to automate the retraining pipeline when new data arrives in BigQuery. Which Vertex AI feature should they use?

A.Vertex AI Prediction
B.Vertex AI Pipelines
C.Vertex AI Model Registry
D.Vertex AI Feature Store
AnswerB

Pipelines can be scheduled or triggered by events to automate ML workflows.

Why this answer

Vertex AI Pipelines is the correct choice because it enables you to define, automate, and orchestrate end-to-end ML workflows, including retraining models when new data arrives. By integrating with BigQuery triggers or Cloud Scheduler, you can set up a pipeline that automatically ingests new data, preprocesses it, retrains the model, and deploys the updated version—all without manual intervention.

Exam trap

CompTIA often tests the distinction between operational tools (like Prediction or Model Registry) and orchestration tools (like Pipelines), so the trap here is confusing a component that manages models or features with the service that actually automates the end-to-end retraining workflow.

How to eliminate wrong answers

Option A is wrong because Vertex AI Prediction is a serving endpoint for deploying models to make predictions, not a tool for automating retraining pipelines. Option C is wrong because Vertex AI Model Registry is a central repository for managing model versions and metadata, but it does not orchestrate the retraining workflow itself. Option D is wrong because Vertex AI Feature Store is designed for managing and serving feature data consistently across training and serving, not for automating pipeline execution.

84
Multi-Selectmedium

A computer vision team is preparing a model for deployment to a fleet of low-power cameras that run on battery and have limited RAM. They want to reduce model size and inference cost while keeping accuracy acceptable for detecting a small set of object classes. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Post-training quantization of weights to 8-bit integers
B.Pruning near-zero weight connections followed by fine-tuning
C.Switching the training optimizer from SGD to AdamW
D.Training with a larger batch size on the same dataset
E.Increasing the number of convolutional filters in each layer
AnswersA, B

Post-training quantization converts 32-bit floating-point weights and activations to 8-bit integers, cutting model size by roughly four times and enabling faster integer arithmetic on constrained hardware. It requires no retraining, so the team can apply it to an existing model quickly. For a small set of object classes, the accuracy loss is usually small enough to remain acceptable on battery-powered cameras.

Why this answer

Model compression for constrained devices typically combines reduced numerical precision with reduced parameter count. Quantizing weights to 8-bit integers shrinks the model and speeds integer inference without retraining, while pruning unimportant connections and fine-tuning recovers accuracy. Both directly lower RAM use and inference energy, which are the binding constraints for battery-powered cameras.

Exam trap

The trap here is confusing training-time hyperparameter changes, such as optimizer or batch size, with deployment-time model compression techniques.

85
MCQeasy

A machine learning engineer needs to train a deep neural network on a large image dataset. Which hardware component is specifically optimized for this task due to its high parallel processing capability and is commonly used in AI training?

A.Central Processing Unit (CPU)
B.Neural Processing Unit (NPU)
C.Graphics Processing Unit (GPU)
D.Tensor Processing Unit (TPU)
AnswerC

GPUs have thousands of cores that excel at parallel processing, making them the industry standard for training deep neural networks.

Why this answer

Graphics Processing Units (GPUs) are specifically optimized for the parallel processing required in deep neural network training. Their architecture contains thousands of smaller cores designed to handle multiple matrix operations simultaneously, which is the core computation in backpropagation and forward passes of neural networks. This makes GPUs the standard choice for training large image datasets in AI.

Exam trap

CompTIA often tests the distinction between training and inference hardware, where candidates may confuse NPUs (optimized for inference) with GPUs (optimized for training), or assume TPUs are the most common due to their specialization, when GPUs remain the industry standard for deep learning training.

How to eliminate wrong answers

Option A is wrong because CPUs are optimized for sequential, low-latency processing with a small number of powerful cores, not the massive parallelism needed for deep learning matrix operations. Option B is wrong because Neural Processing Units (NPUs) are specialized for inference (running trained models) with lower power consumption, not for the heavy parallel training workloads that GPUs handle. Option D is wrong because Tensor Processing Units (TPUs) are custom ASICs designed by Google specifically for TensorFlow workloads, but they are less commonly used in general AI training compared to GPUs, and the question asks for the hardware 'commonly used' in AI training, which is the GPU.

86
MCQhard

An organization must ensure that an AI model deployed on an IoT device meets stringent latency requirements. The model is currently in FP32 and runs at 200ms per inference on the device; the target is 50ms. Which technique will provide the greatest latency reduction with the least accuracy loss?

A.Quantize the model to INT8
B.Apply weight pruning to remove 50% of parameters
C.Switch from TensorFlow Lite to Core ML
D.Distill the model into a smaller architecture
AnswerA

INT8 quantization reduces bit width from 32 to 8, accelerating arithmetic and memory access, often achieving ~4x latency reduction.

Why this answer

Quantizing the model from FP32 to INT8 reduces the precision of weights and activations, which directly decreases memory bandwidth and computational load. On IoT devices with limited resources, this typically yields a 2-4x speedup, bringing the 200ms inference time close to the 50ms target, while INT8 quantization often retains over 90% of the original accuracy when using calibration techniques.

Exam trap

CompTIA often tests the misconception that any optimization technique (like pruning or framework switching) can achieve the same latency reduction as quantization, but only INT8 quantization directly addresses the computational precision bottleneck to deliver the required 4x speedup with minimal accuracy loss.

How to eliminate wrong answers

Option B is wrong because weight pruning removes parameters but does not reduce the precision of the remaining values; the model still operates in FP32, so the latency reduction is limited (often 20-30%) and may not achieve the 4x speedup needed, while aggressive pruning can cause significant accuracy loss. Option C is wrong because switching from TensorFlow Lite to Core ML is a framework change that may optimize for Apple hardware but does not inherently reduce computational precision or model size; it typically provides marginal latency improvements (10-20%) and is platform-specific, not a general solution for the required 4x reduction. Option D is wrong because knowledge distillation creates a smaller student model, but training a new architecture from scratch is time-consuming and may not guarantee the exact 50ms target; the latency reduction depends on the student model's size and hardware compatibility, and distillation often requires extensive retuning to avoid accuracy degradation.

87
MCQhard

An AI research group trains a large language model across a cluster of GPU nodes. They observe that training throughput drops sharply whenever gradient synchronization occurs, and profiling shows GPUs idle while waiting for parameter updates to be exchanged. The model must remain mathematically identical to single-node training. Which change should the team make?

A.Use lower-precision gradient compression that quantizes gradients to 8 bits before transmission.
B.Switch to a distributed data-parallel strategy that overlaps gradient communication with backward computation, such as ring all-reduce with bucketing.
C.Increase the number of gradient accumulation steps so synchronization happens less often.
D.Reduce the global batch size so that fewer gradients need to be exchanged per step.
AnswerB

Ring all-reduce exchanges gradients in chunks around the cluster so bandwidth is used evenly, and bucketing lets reduction of early layers begin while later layers are still computing backward. Communication then overlaps computation instead of serializing after it, cutting GPU idle time while producing the same averaged gradients, so the model remains mathematically identical to single-node training.

Why this answer

The idle time is caused by communication serialized after each backward pass. Ring all-reduce with bucketed, overlapped communication hides most of that transfer behind computation, so GPUs stay busy and the averaged gradient is unchanged. Because the arithmetic of the reduction is preserved, the resulting training run matches single-node behavior numerically.

Exam trap

The trap here is assuming that reducing communication frequency or volume is enough, when the real gain comes from overlapping communication with computation rather than shrinking it.

88
Multi-Selectmedium

A retail chain is deploying an AI-powered demand forecasting system across 500 stores. The system ingests daily sales, weather, and promotion data, and must produce forecasts that update as new data arrives. The MLOps team needs to ensure the deployed model remains accurate over time as consumer behavior shifts. Which TWO practices should they implement? (Choose two.)

Select 2 answers
A.Lock the model version after initial deployment to ensure reproducibility and avoid unexpected behavior changes in production.
B.Schedule regular retraining of the model on a rolling window of recent data, with validation against a holdout set before promotion to production.
C.Increase the model's complexity by adding more layers and parameters to improve its ability to fit historical sales patterns.
D.Set up automated data drift and concept drift monitoring with alerts when statistical properties of input features or prediction errors deviate from training baselines.
E.Deploy the model as a batch job that runs once per quarter to reduce infrastructure costs and operational overhead.
AnswersB, D

Periodic retraining on recent data allows the model to learn current demand patterns, seasonality, and promotion responses. A rolling window keeps the training set relevant while a holdout validation ensures the retrained model outperforms the incumbent before deployment. This directly counters concept drift and maintains forecast accuracy as consumer behavior evolves across the retail chain.

Why this answer

Maintaining forecast accuracy in a dynamic retail environment requires both detecting when the model's assumptions no longer hold and periodically updating the model with recent data. Automated drift monitoring identifies when input distributions or error patterns deviate from training baselines, triggering investigation. Scheduled retraining on a rolling window with validation ensures the model incorporates current trends and promotions.

Together, these practices form a continuous improvement loop that counters concept drift.

Exam trap

The trap here is focusing on model complexity or deployment convenience instead of the operational practices—drift detection and retraining—that actually sustain accuracy as data distributions change.

89
MCQmedium

A company wants to build an AI pipeline that processes streaming data from IoT sensors, performs feature engineering, trains a model incrementally, and deploys the updated model. Which data pipeline technology is BEST suited for the streaming ingestion step?

A.Amazon S3
B.Apache Spark
C.Apache Airflow
D.Apache Kafka
AnswerD

Apache Kafka provides the distributed, partitioned, fault-tolerant publish-subscribe log needed to ingest high-volume streaming sensor data reliably, decoupling producers from the downstream feature engineering and incremental training stages. It satisfies the streaming ingestion constraint better than batch-oriented pipeline technologies.

Why this answer

Apache Kafka is the best choice for the streaming ingestion step because it is a distributed event streaming platform designed for high-throughput, fault-tolerant ingestion of real-time data streams. It acts as a durable message broker that can ingest IoT sensor data in real time and make it available for downstream processing, which aligns perfectly with the requirement for streaming data ingestion.

Exam trap

CompTIA often tests the distinction between data ingestion (Kafka), data processing (Spark), and data storage (S3), so the trap here is confusing Apache Spark's streaming capability with a dedicated ingestion tool, leading candidates to choose Spark instead of Kafka.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is an object storage service designed for batch storage of static files, not for real-time streaming ingestion; it lacks the low-latency publish-subscribe mechanism needed for streaming data. Option B is wrong because Apache Spark is a distributed processing engine that can handle streaming data via Spark Streaming, but it is not a data ingestion technology—it consumes data from sources like Kafka rather than ingesting it directly. Option C is wrong because Apache Airflow is a workflow orchestration tool for scheduling and managing batch pipelines, not a real-time streaming ingestion platform; it cannot handle continuous, low-latency data streams.

90
Multi-Selecthard

A machine learning engineer is designing a pipeline to train a computer vision model using PyTorch on a large dataset stored in an S3 data lake. They need to preprocess images (resize, normalize) and stream them efficiently to GPUs. Which THREE components are essential in this pipeline? (Select THREE.)

Select 3 answers
A.GPU-accelerated training with CUDA
B.CPU-only inference pipeline
C.Apache Airflow to orchestrate the training job
D.PyTorch DataLoader with multi-processing for batching and shuffling
E.Distributed data parallel (DDP) training across multiple GPUs
AnswersA, D, E

GPU acceleration is essential for fast training of deep neural networks.

Why this answer

GPU-accelerated training with CUDA is essential for efficiently training computer vision models on large datasets. PyTorch leverages CUDA to parallelize tensor operations and model computations on NVIDIA GPUs, which is critical for reducing training time from days to hours when processing high-resolution images.

Exam trap

CompTIA often tests the distinction between essential pipeline components (like GPU acceleration and efficient data loading) versus optional orchestration tools (like Airflow) that are not required for the core training loop.

91
MCQeasy

An AI team wants to version control datasets, track experiments, and log model parameters across multiple projects. Which MLOps platform is specifically designed for experiment tracking and model management?

A.MLflow
B.SageMaker Pipelines
C.Vertex AI Pipelines
D.Kubeflow
AnswerA

MLflow directly provides experiment tracking, parameter logging and model registry, matching the stem's requirement to version datasets and log parameters across projects. Its tracking server and model registry components are purpose-built for MLOps workflows, unlike general-purpose version control or CI tooling.

Why this answer

MLflow is an open-source MLOps platform specifically designed for experiment tracking, model management, and reproducibility. It provides a unified API to log parameters, metrics, and artifacts across multiple projects, making it the correct choice for versioning datasets, tracking experiments, and managing models.

Exam trap

CompTIA often tests the distinction between general-purpose pipeline orchestration tools (like SageMaker Pipelines, Vertex AI Pipelines, and Kubeflow) and purpose-built experiment tracking platforms (like MLflow), so the trap is assuming any pipeline tool inherently includes experiment tracking and model management capabilities.

How to eliminate wrong answers

Option B (SageMaker Pipelines) is wrong because it is a fully managed CI/CD service for building, training, and deploying ML pipelines on AWS, but it is not specifically designed for experiment tracking and model management; it focuses on workflow orchestration. Option C (Vertex AI Pipelines) is wrong because it is a serverless ML pipeline service on Google Cloud that orchestrates training and deployment workflows, but it lacks the dedicated experiment tracking and model registry features that MLflow provides. Option D (Kubeflow) is wrong because it is a Kubernetes-native platform for deploying and managing ML workflows, but its primary focus is on orchestration and portability across clusters, not on experiment tracking and model management as a core feature.

92
MCQeasy

Which of the following is a key advantage of using ONNX (Open Neural Network Exchange) format for model deployment?

A.It automatically quantizes models to INT8
B.It enables framework interoperability for model inference
C.It compresses model size by 90%
D.It reduces training time
AnswerB

ONNX defines a common graph and operator format, so a model trained in one framework can be executed by runtimes in another. This framework interoperability for inference is the format's core advantage, decoupling training tooling from deployment runtime choice.

Why this answer

ONNX provides a standardized, open format for representing machine learning models, enabling seamless interoperability between different frameworks (e.g., PyTorch, TensorFlow, scikit-learn). This allows a model trained in one framework to be deployed for inference using a different runtime or hardware accelerator without requiring retraining or manual conversion, which is a key advantage in heterogeneous production environments.

Exam trap

CompTIA often tests the misconception that ONNX provides built-in performance optimizations like quantization or compression, when in fact its primary value is framework interoperability, and any performance gains come from the runtime or additional tools, not the format itself.

How to eliminate wrong answers

Option A is wrong because ONNX does not automatically quantize models to INT8; quantization is a separate optimization step that can be applied to ONNX models using tools like ONNX Runtime or Intel Neural Compressor, but it is not an inherent feature of the format itself. Option C is wrong because ONNX does not inherently compress model size by 90%; while ONNX models may be slightly more compact than some framework-specific formats due to serialization, significant compression requires techniques like pruning or quantization, and 90% reduction is not guaranteed. Option D is wrong because ONNX is a model representation format for inference and interoperability, not a training framework; it does not reduce training time, which depends on the training framework, hardware, and algorithm used.

93
Multi-Selectmedium

A financial institution is deploying an AI model for credit scoring. The model must be explainable to regulators, and the team needs to understand which features contribute most to individual predictions. Which TWO techniques should they use? (Choose two.)

Select 2 answers
A.Principal Component Analysis (PCA)
B.SHAP (SHapley Additive exPlanations)
C.LIME (Local Interpretable Model-agnostic Explanations)
D.Feature importance from a random forest
E.t-SNE (t-Distributed Stochastic Neighbor Embedding)
AnswersB, C

SHAP is a game-theoretic approach that assigns each feature an importance value for a particular prediction. It provides consistent and locally accurate explanations, which are essential for regulatory compliance in credit scoring. SHAP values can be visualized to show how each feature pushes the prediction from the base value.

Why this answer

For explaining individual predictions in a credit scoring model, local explanation techniques are necessary. SHAP and LIME both provide per-instance feature importance, which can be presented to regulators to justify decisions. SHAP offers a solid theoretical foundation, while LIME is flexible and model-agnostic.

Together, they cover the need for explainability.

Exam trap

The trap here is confusing global feature importance methods with local explanation techniques, which are required for individual prediction transparency.

94
MCQhard

A media company wants to automatically generate concise summaries of lengthy earnings-call transcripts. The transcripts average 45 minutes of speech and contain domain-specific financial terminology. The team needs a solution that captures long-range dependencies and produces fluent, abstractive summaries without training a model from scratch. Which approach is most appropriate?

A.Fine-tune a pretrained encoder-decoder transformer (e.g., BART or T5) on the earnings-call corpus using abstractive summarization objectives.
B.Apply latent Dirichlet allocation (LDA) topic modeling to identify key themes, then generate summaries from the top topics.
C.Use an extractive summarization algorithm like TextRank to select the most important sentences from each transcript.
D.Train a large LSTM-based sequence-to-sequence model from scratch on the transcripts with attention.
AnswerA

Encoder-decoder transformers pretrained on large text corpora already understand language structure and can be fine-tuned on domain transcripts to learn financial terminology and summary style. BART and T5 are specifically designed for sequence-to-sequence tasks like abstractive summarization, and their self-attention captures long-range dependencies across thousands of tokens. Fine-tuning avoids training from scratch while adapting to the domain, making this the most effective and efficient choice.

Why this answer

Abstractive summarization of long, domain-specific transcripts is best served by fine-tuning a pretrained encoder-decoder transformer. Models like BART and T5 are pretrained on vast text corpora, giving them strong language understanding, and their architecture handles long-range dependencies through self-attention. Fine-tuning on earnings-call data adapts them to financial terminology and summary style without the prohibitive cost of training from scratch.

This balances quality, fluency, and practicality.

Exam trap

The trap here is assuming extractive methods or topic models can satisfy a requirement for fluent abstractive summaries, when only generative sequence-to-sequence models actually paraphrase and synthesize content.

95
MCQeasy

Which open-source framework is commonly used for building, training, and deploying machine learning models and provides high-level APIs like Keras?

A.TensorFlow
B.Hugging Face Transformers
C.scikit-learn
D.PyTorch
AnswerA

TensorFlow provides Keras and is widely used for production ML.

Why this answer

TensorFlow is the correct answer because it is the open-source framework that provides high-level APIs like Keras for building, training, and deploying machine learning models. Keras, now integrated as tf.keras, offers a user-friendly interface for rapid prototyping while TensorFlow handles the underlying computation graph, distributed training, and model serving via TensorFlow Serving.

Exam trap

Candidates often confuse PyTorch as the only framework with dynamic computation graphs and high-level APIs, but the question specifically asks for the framework that provides Keras, which is exclusive to TensorFlow.

How to eliminate wrong answers

Option B (Hugging Face Transformers) is wrong because it is a specialized library for natural language processing (NLP) models like BERT and GPT, not a general-purpose framework for building and deploying any ML model, and it does not natively include Keras as its high-level API. Option C (scikit-learn) is wrong because it is designed for traditional machine learning algorithms (e.g., decision trees, SVMs) and lacks deep learning capabilities, GPU acceleration, and a high-level API like Keras for neural networks. Option D (PyTorch) is wrong because, although it is a popular deep learning framework, it does not provide Keras as its high-level API; instead, it uses torch.nn and higher-level wrappers like Lightning or Fastai, and Keras is specifically integrated with TensorFlow.

96
MCQhard

A team is deploying a machine learning model on a Kubernetes cluster. They need to ensure low-latency inference and efficient resource utilization. Which approach should they use to dynamically scale inference pods based on request volume?

A.Use a Job resource to process requests in batch
B.Deploy a single large pod on a powerful node
C.Use a Horizontal Pod Autoscaler (HPA) with target CPU utilization
D.Set a fixed number of pod replicas equal to the maximum expected load
AnswerC

Horizontal Pod Autoscaler adjusts replica counts from observed CPU utilisation, satisfying the low-latency and efficient-resource constraint by matching pod capacity to request-driven load. It scales horizontally within the cluster, so inference pods expand as traffic rises and contract when idle, avoiding the over-provisioning that fixed replicas would cause.

Why this answer

The Horizontal Pod Autoscaler (HPA) is the correct choice because it automatically scales the number of inference pods based on observed CPU utilization or custom metrics, ensuring low-latency inference by adding replicas during traffic spikes and reducing waste during idle periods. This dynamic scaling aligns with the need for efficient resource utilization in a Kubernetes cluster, as it adjusts pod count in real-time to match request volume without manual intervention.

Exam trap

A common misconception is that batch processing (Jobs) or static scaling is suitable for real-time inference, when in fact dynamic scaling with HPA is required to balance latency and resource efficiency in Kubernetes.

How to eliminate wrong answers

Option A is wrong because a Job resource is designed for batch processing and runs pods to completion, not for serving continuous inference requests that require low-latency responses; it cannot dynamically scale based on request volume. Option B is wrong because deploying a single large pod on a powerful node creates a single point of failure and cannot handle variable request loads efficiently, leading to either over-provisioning or under-provisioning and increased latency during spikes. Option D is wrong because setting a fixed number of pod replicas equal to the maximum expected load wastes resources during low-traffic periods and fails to adapt to actual request volume, contradicting the goal of efficient resource utilization.

97
MCQmedium

A company needs to store large volumes of unstructured data (PDFs, images, logs) for future AI model training. The data must be easily accessible by data scientists using Spark and must support cost-effective storage. Which data infrastructure is MOST appropriate?

A.Snowflake data warehouse
B.Relational database like Amazon RDS
C.Pinecone vector database
D.Amazon S3 data lake
AnswerD

Amazon S3 provides durable, cost-effective object storage for unstructured PDFs, images and logs, and integrates natively with Spark and analytics tooling. This satisfies both the accessibility requirement for data scientists and the cost-effective storage constraint for future AI training.

Why this answer

Amazon S3 is the canonical data lake storage layer for large volumes of unstructured data such as PDFs, images, and logs, and it integrates natively with Spark via the S3A connector and with AWS Glue, EMR, and Athena. Its object storage model, tiered storage classes (Standard, IA, Glacier), and pay-for-what-you-use pricing make it cost-effective for petabyte-scale AI training corpora. This combination of scalability, accessibility, and cost is exactly what the scenario requires.

Exam trap

AI0-001 often tests the confusion between storage layers (S3 data lake) and compute/query layers (Snowflake, RDS) or specialized stores (Pinecone), so candidates who focus on 'analytics' rather than 'unstructured storage' pick the wrong tier.

How to eliminate wrong answers

Option A is wrong because Snowflake is a structured/semi-structured analytical data warehouse optimized for SQL workloads, not a cost-effective store for raw unstructured PDFs and images. Option B is wrong because Amazon RDS is a relational database designed for transactional structured data with row/column schemas, and it cannot economically or practically store large binary objects at petabyte scale. Option C is wrong because Pinecone is a vector database for storing embeddings and performing similarity search, not a general-purpose object store for raw unstructured files.

98
MCQhard

An AI platform team is building a retrieval-augmented generation service over an internal knowledge base of roughly 40 million technical documents. Queries must return semantically relevant passages in under 50 ms at the vector search layer. The team wants approximate nearest neighbor search that supports metadata filtering on fields such as product line and document date, and they want to avoid a separate relational database for those filters. Which vector index type best matches these requirements?

A.HNSW index with payload filtering
B.Flat (brute-force) index
C.Product quantization index without a graph
D.Inverted file index with a very large nlist
AnswerA

HNSW builds a hierarchical navigable small-world graph that delivers high recall at very low latency on tens of millions of vectors, and modern engines support payload or metadata filtering combined with the graph traversal. That satisfies both the sub-50 ms semantic search target and the product-line and date constraints without adding a separate relational store for filters.

Why this answer

The combination of tens of millions of vectors, a strict latency ceiling, and integrated metadata filtering points to a graph-based approximate index with payload filtering. HNSW's layered graph gives logarithmic-style search with high recall, and payload filtering lets the engine apply product-line and date constraints during traversal rather than post-filtering, which would otherwise shrink the candidate pool and hurt recall.

Exam trap

The trap here is treating metadata filtering as something applied after vector search, when post-filtering at this scale can silently drop most candidates and destroy recall.

99
MCQeasy

Which hardware accelerator is specifically designed by Google for training and inference of machine learning models, particularly their TensorFlow framework?

A.NPU
B.FPGA
C.GPU
D.TPU
AnswerD

Tensor Processing Units are Google-designed ASICs built around systolic arrays for matrix multiplication, purpose-built for TensorFlow training and inference workloads. Unlike GPUs, which are general-purpose parallel processors, TPUs deliver higher throughput per watt specifically for the tensor operations the stem describes.

Why this answer

TPU (Tensor Processing Unit) is Google's custom ASIC designed to accelerate ML workloads, especially with TensorFlow.

100
MCQhard

An ML team uses Kubeflow to orchestrate a pipeline that includes data preprocessing, model training, and evaluation. The pipeline runs on a Kubernetes cluster. After a cluster upgrade, the pipeline fails at the training step with an 'OOMKilled' error. What is the MOST likely cause?

A.The training code has a memory leak
B.The pipeline definition is missing a step dependency
C.The Kubernetes node's memory resources were not correctly allocated to the pod's resource requests or limits
D.The training data is corrupted
AnswerC

OOMKilled means the container exceeded its memory limit and the kernel terminated it. If the pod's memory requests or limits were not adjusted for the upgraded node's capacity, the training container gets killed, so misallocated pod memory resources are the most likely cause.

Why this answer

Kubeflow pipelines run each step as a Kubernetes pod, and the OOMKilled status is a Kubernetes-level signal that the container exceeded its memory limit (or the node ran out of allocatable memory). After a cluster upgrade, node instance types, kubelet reservations, or resource quotas often change, so the pod's original requests/limits no longer match available memory. The error appears specifically at the training step because that step has the largest memory footprint, making it the first to be evicted or killed.

Exam trap

The trap here is assuming OOMKilled points to application code (a leak) rather than to Kubernetes resource configuration, when the timing after a cluster upgrade is the decisive clue.

How to eliminate wrong answers

Option A is wrong because a code memory leak would typically cause gradual memory growth and would not correlate with a cluster upgrade; it also would not produce an immediate OOMKilled at the same pipeline stage unless limits were already too tight. Option B is wrong because a missing step dependency causes the pipeline to fail with a DAG/ordering error or a step running before its inputs exist, not an OOMKilled termination. Option D is wrong because corrupted training data produces data parsing, schema, or accuracy errors, not a container memory kill signal from the kubelet.

101
MCQmedium

A data engineer is building a pipeline to process streaming clickstream data and feed it into a real-time ML feature store. Which tool is BEST suited for the streaming ingestion?

A.Amazon S3
B.Apache Airflow
C.Apache Spark (batch mode)
D.Apache Kafka
AnswerD

Apache Kafka provides a distributed, partitioned commit log with durable ordered ingestion and replay, letting streaming clickstream events feed a real-time feature store with low latency. Batch-oriented tools cannot satisfy the continuous, real-time ingestion requirement.

Why this answer

Apache Kafka is the best tool for streaming ingestion because it is a distributed event streaming platform designed for high-throughput, low-latency ingestion of real-time data streams. It can handle clickstream data and feed it into a feature store with minimal delay, supporting exactly-once semantics and scalability.

Exam trap

AI0-001 often tests the distinction between batch and streaming tools, and candidates may incorrectly choose Airflow or Spark for real-time ingestion due to familiarity.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is object storage designed for batch processing, not real-time streaming ingestion; it has higher latency and is not optimized for continuous data streams. Option B is wrong because Apache Airflow is a workflow orchestration tool for batch pipelines, not a streaming ingestion system; it schedules tasks but does not handle real-time data flow. Option C is wrong because Apache Spark in batch mode processes data in discrete chunks, not continuously; while Spark Streaming exists, the option specifies batch mode, which is unsuitable for real-time ingestion.

102
Multi-Selecthard

A company is building a secure AI system that must comply with GDPR. They want to allow users to request deletion of their personal data from training sets and model outputs. Which THREE techniques should they implement?

Select 3 answers
A.Model ensembling
B.Differential privacy
C.Data retention and deletion policies
D.Machine unlearning
E.Federated learning
AnswersB, C, D

Differential privacy ensures that the model does not memorize individual data points.

Why this answer

Differential privacy (B) is correct because it adds calibrated noise to training data or model outputs, ensuring that the inclusion or exclusion of any individual's data does not significantly affect the model's behavior. This provides a mathematical guarantee of privacy, which is essential for GDPR compliance when handling personal data. By limiting information leakage, differential privacy helps protect user data even if deletion requests are not fully implemented.

Exam trap

CompTIA AI often tests the misconception that federated learning alone satisfies GDPR deletion requirements, when in fact it only addresses data locality, not the ability to remove a specific user's influence from a trained model.

103
MCQmedium

A team is building a retrieval-augmented generation (RAG) pipeline. They need to store embeddings of company documents and perform fast similarity searches. Which data store is BEST suited for this task?

A.Snowflake
B.Pinecone
C.Apache Kafka
D.Amazon S3
AnswerB

Pinecone is a purpose-built vector database supporting approximate nearest-neighbour indexes over high-dimensional embeddings, delivering the fast similarity search the RAG pipeline requires. Relational stores lack native vector indexing, so they cannot meet the low-latency retrieval constraint for document embeddings.

Why this answer

Pinecone is a purpose-built vector database designed for storing and querying high-dimensional embeddings with fast approximate nearest neighbor (ANN) search. In a RAG pipeline, embeddings of company documents must be retrieved quickly to feed relevant context to the LLM, and Pinecone’s optimized indexing (e.g., HNSW or IVF) and serverless scaling make it the ideal choice for this task.

Exam trap

The trap here is that candidates may confuse general-purpose storage (like S3 or Snowflake) with specialized vector databases, assuming any database can handle embeddings efficiently, but CompTIA AI tests the understanding that only purpose-built vector stores provide the required ANN search performance for RAG.

How to eliminate wrong answers

Option A is wrong because Snowflake is a cloud data warehouse optimized for SQL-based analytical queries on structured data, not for low-latency vector similarity searches on embeddings. Option C is wrong because Apache Kafka is a distributed event streaming platform for real-time data pipelines and message brokering, not a storage and retrieval system for vector embeddings. Option D is wrong because Amazon S3 is an object storage service for static files and does not natively support vector indexing or similarity search operations.

104
MCQeasy

A developer is building a mobile app that uses a pre-trained image classification model on-device. Which framework should they use to run the model on iOS devices?

A.Hugging Face Transformers
B.TensorFlow Lite
C.PyTorch Mobile
D.Core ML
AnswerD

Core ML is Apple's native on-device inference framework, so it runs the pre-trained image classification model directly on iOS hardware without a network round trip. It satisfies the stem's on-device constraint, unlike cloud-hosted alternatives, and integrates with Xcode tooling to convert and optimise models for Apple silicon.

Why this answer

Core ML is the correct framework because it is Apple's machine learning framework designed specifically for on-device inference on iOS, macOS, watchOS, and tvOS. It provides optimized performance and integration with Apple's hardware, making it the best choice for running a pre-trained image classification model on iOS devices.

Exam trap

AI0-001 often tests the choice of framework for specific platforms, and candidates may choose cross-platform tools like TensorFlow Lite when a native solution like Core ML is more appropriate.

How to eliminate wrong answers

Option A is wrong because Hugging Face Transformers is a library for natural language processing, not optimized for on-device image classification on iOS. Option B is wrong because TensorFlow Lite is a framework for on-device inference but is cross-platform; while it can run on iOS, Core ML is more native and optimized for Apple devices. Option C is wrong because PyTorch Mobile is also cross-platform and can run on iOS, but Core ML is the preferred and most integrated solution for iOS.

105
MCQhard

A logistics company runs a route-optimization model on a Kubernetes cluster. During peak hours the inference pods are frequently evicted because the nodes run out of memory, even though average GPU utilization stays below 40 percent. The team wants to reduce evictions without changing the model or adding nodes. Which action best addresses the root cause?

A.Lower the container image size by switching to a distroless base image for the inference service.
B.Enable horizontal pod autoscaling on GPU utilization so additional replicas start during peak hours.
C.Set explicit CPU and memory requests and limits on the inference containers so the scheduler can place and protect them accurately.
D.Increase the GPU memory allocation per pod by requesting additional nvidia.com/gpu resources.
AnswerC

Evictions driven by node memory pressure usually mean the pods have no memory requests, so the scheduler overcommits the node and the kubelet later reclaims memory by evicting workloads. Declaring realistic requests lets the scheduler reserve capacity and declaring limits bounds each pod, which stops one inference process from consuming memory that other pods depend on and prevents pressure-driven eviction.

Why this answer

Memory-pressure eviction happens when the kubelet must reclaim node memory, and pods without declared memory requests are the first to be sacrificed because the scheduler never reserved capacity for them. Setting realistic requests and limits makes placement accurate and bounds each pod's consumption, so the node stays below its eviction threshold without changing the model or adding hardware.

Exam trap

The trap here is chasing GPU symptoms when the eviction signal points to host memory pressure, so adding accelerators or replicas leaves the actual cause untouched.

106
MCQeasy

A startup is training a recommendation model on a single workstation with one GPU. The dataset has grown to 2 TB, and training now takes several days. The team wants to reduce training time by adding more GPUs to the same workstation. Which technology should they use to enable efficient multi-GPU training with minimal code changes?

A.Apache Hadoop MapReduce
B.CUDA Multi-Process Service (MPS)
C.NVIDIA NCCL with DistributedDataParallel
D.TensorFlow Lite
AnswerC

NCCL provides optimized inter-GPU communication, and DistributedDataParallel in PyTorch enables data-parallel training across multiple GPUs with minimal code changes. This combination is the standard for scaling training on a single node with multiple GPUs, directly reducing training time.

Why this answer

NCCL with DistributedDataParallel is the standard approach for multi-GPU training within a single node. It provides efficient gradient synchronization and requires minimal code changes in PyTorch. The other options are either for data processing, single-GPU sharing, or edge inference, none of which solve the training time issue.

Exam trap

The trap here is assuming that any parallel computing framework, such as Hadoop, can accelerate GPU training, when it lacks the necessary communication primitives.

107
MCQmedium

A computer vision team trains a convolutional neural network for manufacturing defect detection on a workstation with an NVIDIA RTX A6000 GPU. They want to reduce training time by increasing throughput without changing model architecture or batch size. Which action should they take?

A.Increase the batch size to the maximum that fits in GPU memory and keep the learning rate unchanged.
B.Shard the dataset across multiple CPU cores using NumPy and disable GPU acceleration for the convolution layers.
C.Enable mixed-precision training using NVIDIA Tensor Cores with FP16 compute and FP32 accumulation.
D.Convert the trained model to TensorFlow Lite and redeploy it for training on the GPU.
AnswerC

Mixed-precision training uses Tensor Cores to perform matrix multiplications in FP16 while accumulating in FP32, roughly doubling throughput on Ampere-class GPUs without altering the model architecture or batch size. It preserves numerical stability because accumulation stays in FP32. This directly reduces training time for the existing CNN on the RTX A6000.

Why this answer

Mixed-precision training exploits Tensor Cores to compute in FP16 while accumulating in FP32, delivering substantial throughput gains on modern NVIDIA GPUs with minimal accuracy impact. Because the architecture and batch size remain unchanged, it fits the constraint precisely. Other listed actions either alter training semantics or move work to slower hardware, so they do not satisfy the goal.

Exam trap

The trap here is assuming that any change to precision degrades model accuracy, when mixed-precision keeps FP32 accumulation specifically to preserve numerical stability.

108
MCQeasy

Which AI accelerator is specifically designed by Google to accelerate the training and inference of large neural networks, especially in their cloud environment?

A.GPU
B.NPU
C.TPU
D.FPGA
AnswerC

Google-designed TPUs are application-specific integrated circuits built around systolic array matrix multiplication, which suits the massive parallel tensor operations in neural network training and inference. This directly satisfies the stem's requirement for a Google-designed accelerator operating within Google Cloud, unlike GPUs or general-purpose CPUs.

Why this answer

The Tensor Processing Unit (TPU) is Google's custom-designed ASIC specifically built to accelerate the training and inference of large neural networks. Unlike general-purpose hardware, TPUs are optimized for TensorFlow workloads and are a core component of Google Cloud's AI infrastructure, offering high throughput for matrix operations common in deep learning.

Exam trap

CompTIA often tests the distinction between custom-designed accelerators (like TPU) and general-purpose or reconfigurable hardware (like GPU, NPU, FPGA), expecting candidates to know that TPU is Google's proprietary solution for neural network acceleration in their cloud.

How to eliminate wrong answers

Option A is wrong because GPUs (Graphics Processing Units) are general-purpose parallel processors designed for graphics and compute, not specifically by Google for neural network acceleration in their cloud; they are widely used but not Google's custom accelerator. Option B is wrong because NPU (Neural Processing Unit) is a generic term for processors designed to accelerate neural networks, but it is not a specific Google-designed chip; Google's custom accelerator is the TPU. Option D is wrong because FPGAs (Field-Programmable Gate Arrays) are reconfigurable hardware that can be programmed for various tasks, but they are not specifically designed by Google for neural network training and inference in their cloud environment; Google uses TPUs for that purpose.

109
Multi-Selecthard

An organisation is deploying a fine-tuned LLM for internal use. They need to ensure the API endpoint is secure and cost-effective. Which TWO measures should they implement? (Choose 2)

Select 2 answers
A.Implement API key authentication
B.Enable content filtering
C.Disable logging to reduce storage costs
D.Apply rate limiting per user
E.Use gRPC instead of REST
AnswersA, D

API keys restrict access to authorised clients.

Why this answer

API key authentication (Option A) is a fundamental security measure that ensures only authorized clients can access the LLM endpoint. It provides a simple, lightweight mechanism to validate requests without the overhead of full OAuth, making it both secure and cost-effective for internal deployments.

Exam trap

The CompTIA AI+ exam tests the distinction between security measures (authentication, rate limiting) and non-security features (content filtering, protocol choice), leading candidates to mistakenly select content filtering or gRPC as security controls.

110
MCQhard

An MLOps engineer is deploying a scikit-learn random forest model to a Kubernetes cluster for a low-traffic internal API. The team wants to avoid maintaining a custom Flask wrapper and prefers a standard serving solution that supports REST and gRPC. Which serving component should they choose?

A.KServe with a scikit-learn model server runtime.
B.NVIDIA Triton Inference Server with a Python backend script that loads the pickle file.
C.TensorFlow Serving configured with a SavedModel export of the random forest.
D.A Kubernetes Deployment running a Flask app that loads the model and exposes only a REST endpoint.
AnswerA

KServe provides standardized model serving on Kubernetes with built-in support for REST and gRPC, autoscaling, and canary rollouts. Its scikit-learn runtime loads the pickled model directly without a custom Flask wrapper. This matches the requirement for a standard, low-maintenance serving solution for a random forest model on Kubernetes.

Why this answer

KServe offers a purpose-built scikit-learn runtime that serves pickled models over REST and gRPC on Kubernetes, with autoscaling and rollout features included. It eliminates the need for a bespoke Flask wrapper, which the team explicitly wants to avoid. The other options either require custom code, target GPU deep learning workloads, or demand unsupported format conversion.

Exam trap

The trap here is equating 'model serving on Kubernetes' with 'write a Flask container', ignoring that KServe provides ready-made runtimes for common frameworks like scikit-learn.

111
MCQhard

A platform team is preparing a feature store for a recommendation system. They need point-in-time correct feature retrieval so that training datasets do not leak future information, and they need the same features served online with low latency. Which architecture best satisfies both requirements?

A.Use a single relational database table that overwrites feature values in place with no timestamps.
B.Compute features on the fly in the training pipeline and again independently in the serving path using separate code.
C.Maintain an offline store with time-stamped feature values for training and a synchronized online store keyed by entity for low-latency serving.
D.Store features only in a data lake and have the online service query the lake on each request.
AnswerC

A dual-store feature platform keeps historical, time-stamped values in an offline store for point-in-time correct training joins, and materializes the latest values into a low-latency online store keyed by entity ID. This prevents label leakage during training while enabling fast online retrieval. It is the standard architecture for consistent offline and online features.

Why this answer

A feature store that pairs a time-stamped offline store with a synchronized online store supports point-in-time correct training joins and fast online retrieval from the same feature definitions. This dual-store pattern prevents label leakage and reduces training-serving skew. Single-store, lake-query, and duplicated-code approaches each fail at least one of correctness or latency.

Exam trap

The trap here is assuming that one storage system can serve both batch training and low-latency online inference, when the two access patterns require different stores synchronized through the same transformation logic.

112
MCQhard

A data science team is deploying a real-time fraud detection model on edge devices in retail stores. The model must infer under 10ms and fit within 50MB memory. Which combination of techniques should the team apply?

A.Model parallelism and distributed inference
B.Increase batch size and use FP16 precision
C.Train a larger model and use distillation to transfer knowledge
D.Model quantization to INT8 and pruning of low-weight connections
AnswerD

Quantisation to INT8 shrinks weights from 32-bit to 8-bit, cutting memory roughly fourfold, while pruning removes low-weight connections to reduce computation. Together they meet the 50MB footprint and sub-10ms inference latency constraints on constrained edge hardware.

Why this answer

INT8 quantization reduces each weight from 32-bit float to 8-bit integer, cutting model size ~4x and enabling faster integer arithmetic that meets the sub-10ms latency target. Pruning removes low-magnitude weight connections, further shrinking the 50MB footprint and reducing compute. Together they are the standard edge-optimization pairing for latency- and memory-constrained inference.

Exam trap

AI0-001 often tests the misconception that more hardware parallelism or larger batches solve latency problems, when edge constraints actually demand model compression techniques like quantization and pruning.

How to eliminate wrong answers

Option A is wrong because model parallelism splits a model across multiple devices and distributed inference adds network hops — both increase latency and assume multiple nodes, which is impractical on a single constrained edge device. Option B is wrong because increasing batch size raises per-inference latency and memory, and FP16 alone only halves precision without the 4x reduction INT8 provides, so it won't reliably fit 50MB. Option C is wrong because training a larger model then distilling still yields a model that must be quantized/pruned to meet the constraints; distillation alone doesn't guarantee the 50MB/10ms targets.

113
MCQmedium

A financial institution is deploying an AI model that predicts loan default risk. The model is trained on historical data that includes sensitive attributes like zip code and marital status. The compliance team is concerned about disparate impact. Which technique should be applied during model training to mitigate bias while maintaining predictive performance?

A.Apply adversarial debiasing during training to minimize the model's ability to predict the sensitive attribute.
B.Post-process predictions by adjusting thresholds for different demographic groups.
C.Remove all sensitive attributes from the training dataset.
D.Oversample the minority group in the training data to balance representation.
AnswerA

Adversarial debiasing trains the main model to make accurate predictions while simultaneously training an adversary to predict the sensitive attribute from the model's outputs. By minimizing the adversary's success, the model learns representations that are less informative about sensitive attributes, reducing disparate impact. This technique can be integrated into training and often preserves predictive performance better than simply removing sensitive features.

Why this answer

Adversarial debiasing is a training-time technique that reduces the model's reliance on sensitive attributes by pitting the predictor against an adversary. This mitigates disparate impact while often maintaining accuracy. Removing sensitive attributes is insufficient due to proxies, oversampling addresses class imbalance rather than bias, and post-processing occurs after training, which does not meet the scenario's requirement for an in-training method.

Exam trap

The trap here is assuming that removing sensitive attributes guarantees fairness, when proxy variables can perpetuate bias and more robust training-time techniques are needed.

114
MCQeasy

A startup is developing an AI chatbot and wants to use a pre-trained language model to generate responses. They need to integrate the model into their application with minimal latency and cost. Which approach should they take?

A.Use a rule-based chatbot system
B.Train a custom transformer model from scratch on their own data
C.Use a pre-trained model via a managed API service
D.Deploy the model on a local server with a single GPU
AnswerC

Using a pre-trained model through a managed API service, such as OpenAI's GPT or Google's Vertex AI, allows the startup to integrate advanced language capabilities without managing infrastructure. This approach minimizes latency by leveraging the provider's optimized serving and reduces cost by paying only for usage, making it ideal for a startup.

Why this answer

The startup needs to integrate a pre-trained language model with minimal latency and cost. Using a managed API service provides immediate access to state-of-the-art models without infrastructure management, and the pay-as-you-go model keeps costs low. This is the most efficient approach for a startup with limited resources.

Exam trap

The trap here is overlooking the operational and cost benefits of managed API services and instead opting for self-hosted solutions that require significant upfront investment.

← PreviousPage 2 of 2 · 114 questions total

Ready to test yourself?

Try a timed practice session using only AI Infrastructure and Technologies questions.