Courseiva

AI0-001 · domain

AI Infrastructure and Technologies

This domain covers the compute, storage, networking, and pipeline infrastructure that hosts AI workloads. Expect scenario questions on choosing accelerators, deployment topologies, streaming data tooling, and model packaging for cloud inference. You must match a stated constraint — latency, compliance, scale, or existing stack — to the correct service or architecture rather than the most familiar one.

114 questions27 easy51 medium36 hard

Focused practice

Practice AI Infrastructure and Technologies questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about AI Infrastructure and Technologies

Be able to select accelerators, deployment architectures, streaming tools, and model packaging formats from stated constraints. The single most important skill is mapping the constraint — latency, compliance, cloud vendor, or data velocity — to the correct infrastructure choice, not the most familiar one.

Google Cloud TPUs purpose-built for large neural network training and inference

On-premises GPU inference serving for compliance-bound low-latency LLM deployment

Apache Kafka plus stream processing for real-time feature engineering into a data lake

Packaging TensorFlow models as SageMaker model artifacts for scalable AWS inference

Watch out for

Common AI Infrastructure and Technologies exam traps

  • ▸Assuming any GPU instance fits every inference job; ignoring that TPUs are Google-specific and not portable to AWS or on-prem
  • ▸Choosing a batch or serverless pipeline for real-time streaming ML when the requirement is continuous low-latency processing
  • ▸Uploading raw training checkpoints to SageMaker without the expected model artifact structure and inference handler

Question index

All AI Infrastructure and Technologies questions (114)

Click any question to see the full explanation, or start a practice session above.

1

A developer is using Hugging Face Transformers to fine-tune a BERT model for sentiment analysis. They want to track experiments, log metrics, and compare runs. Which MLOps tool should they integrate?

Easy
2

An ML team deploys a model on edge devices using INT8 quantization. They notice a significant drop in accuracy on a subset of classes. Which technique should they apply to recover accuracy without increasing model size?

Hard
3

A company wants to build a real-time anomaly detection system for IoT sensor data using edge AI. The model must run on resource-constrained devices with minimal power consumption. Which model optimization technique is MOST important?

Easy
4

A startup is building a retrieval-augmented generation (RAG) application that must answer questions over a 500,000-document internal knowledge base with low query latency. They plan to use a vector database. Which TWO design choices best support fast, scalable similarity search? (Choose two.)

Medium
5

A data scientist is choosing a hardware accelerator for training a large transformer model. Which of the following is specifically designed for deep learning workloads and offers the highest throughput for matrix multiplications?

Easy
6

A retail company wants to add natural language search to its product catalog. The team plans to convert product descriptions and customer queries into embeddings so that semantically similar items surface even when the wording differs. They need an embedding model that maps text into a dense vector space where cosine similarity reflects meaning. Which type of model should they use?

Easy
7

An ML engineer wants to deploy a model as a REST API that can scale to handle thousands of inference requests per second. Which serving approach is most appropriate?

Easy
8

A machine learning team is training a large transformer model on a text corpus. They need to reduce training time while maintaining model accuracy. Which hardware configuration would be MOST effective for this task?

Medium
9

A team is selecting a vector database for a RAG application that requires low-latency similarity search on millions of embeddings. They prioritize ease of use and fully managed cloud service. Which TWO options meet these requirements?

Medium
10

A healthcare AI team is training a model to predict patient readmission risk from electronic health records. The dataset contains sensitive patient data and must comply with HIPAA. They need to ensure that the model training process does not expose protected health information (PHI) and that the model does not memorize individual patient data. Which technique should they implement?

Hard
11

A company is building a multi-modal AI application that processes text, images, and audio. They need a unified platform to store embeddings for all modalities, perform hybrid search (vector + metadata filtering), and scale to millions of vectors. Which THREE services are suitable for this purpose? (Choose THREE.)

Hard
12

A team wants to deploy a large language model on edge devices with limited memory and compute. They need to reduce model size by at least 50% while preserving accuracy. Which combination of techniques is most effective?

Hard
13

A data science team is deploying a deep learning model for real-time inference on edge devices with limited power and memory. Which model optimisation technique would be MOST effective for reducing latency and memory footprint while maintaining acceptable accuracy?

Medium
14

A data scientist needs to deploy a PyTorch model to production with low-latency inference. The model must be served as a REST API and should support GPU acceleration. Which combination of tools is MOST suitable for this task?

Medium
15

A company uses Azure OpenAI to generate customer support responses. The team notices that repeated queries with similar context incur high costs due to token usage. They want to reduce costs without affecting response quality. Which strategy is MOST effective?

Hard
16

An AI developer needs to store large amounts of unstructured data (e.g., images, logs) for training datasets. Which cloud storage solution is purpose-built for data lakes?

Easy
17

A data scientist is building a recommendation system using Apache Spark for feature engineering. They need to process streaming user click data in real-time before feeding into the model. Which tool should they use for the streaming data ingestion?

Medium
18

A cybersecurity firm is building an anomaly detection system for network traffic. The dataset contains millions of connection records with dozens of features, but only 0.1% are labeled as malicious. The team needs a model that can flag suspicious connections while minimizing false positives that overwhelm analysts. Which approach is most appropriate?

Hard
19

A team is deploying a model on Kubernetes using Kubeflow. They want to automatically scale the number of inference pods based on request latency. Which Kubernetes-native feature should they configure?

Hard
20

A company is building an AI-powered document processing system that extracts information from scanned PDFs. The system must handle varying document layouts and languages. The team wants to use a pre-trained model and fine-tune it on their own data. Which TWO techniques are most appropriate to improve the model's ability to generalize to new document layouts? (Choose two.)

Medium
21

An organization is building a recommendation system that requires low-latency vector similarity search. They need to store and query millions of embeddings. Which THREE technologies are appropriate for this task?

Medium
22

A financial institution requires that all AI model predictions be explainable and auditable for regulatory compliance. Which model serving approach should be used to meet these requirements?

Medium
23

A company uses a cloud-based ML platform to train a model and wants to deploy it for real-time inference. They also need to monitor the endpoint for data drift and retrain automatically. Which feature enables this automated retraining pipeline?

Medium
24

A retail analytics team is choosing a vector database to power semantic search over millions of product descriptions and support retrieval-augmented generation. The team must keep infrastructure costs predictable and needs fast approximate nearest neighbor queries as the index grows. Which TWO characteristics of approximate nearest neighbor indexing should the team evaluate when selecting the vector database? (Choose two.)

Hard
25

A company is implementing a retrieval-augmented generation (RAG) pipeline using a vector database. They notice that the retrieved documents often lack relevance to the query. Which adjustment would MOST improve retrieval quality?

Hard
26

An MLOps team wants to deploy a trained PyTorch model to production with low latency inference. The model must be interoperable across different frameworks and runtimes. Which approach is BEST?

Medium
27

A hospital wants to run a patient-triage natural language model entirely inside its own data center because patient records cannot leave the premises. The IT team needs an inference serving component that exposes an HTTP endpoint, supports model versioning, and can be operated without a managed cloud service. Which technology should the team deploy?

Easy
28

A startup is developing a voice assistant that runs on smart speakers with limited processing power and memory. The team wants to use a pre-trained speech recognition model but needs to reduce its size and latency. Which approach is most suitable?

Easy
29

A data scientist is training a large language model on a custom dataset using PyTorch on AWS. The training is taking too long due to GPU memory constraints. The team wants to use multiple GPUs across instances with minimal code changes. Which AWS service should they use?

Hard
30

A healthcare startup needs to deploy an AI model for real-time patient monitoring on IoT devices with limited battery and compute. The model must run locally with minimal latency. Which TWO strategies are most appropriate?

Medium
31

An organization is deploying a large language model on-premises for compliance reasons. They need to serve inference requests with low latency. Which architecture should they use?

Medium
32

A team is using an API from a cloud AI service to generate text. They notice that repeated requests with the same prompt return different outputs. They want consistent responses for testing. Which parameter should they adjust?

Medium
33

A company is deploying a real-time object detection model on a fleet of IoT cameras. The model must run at 30 FPS on a device with limited memory and no internet connectivity. Which combination of techniques is MOST suitable?

Hard
34

An AI platform team is building a feature store that feeds both offline training jobs and an online model that must return features within a few milliseconds. They are concerned that a feature computed one way during training could be computed differently at serving time. Which design choice best prevents this training-serving skew?

Medium
35

An MLOps engineer is building a feature pipeline for a recommendation model. Features must be served to the online model with single-digit millisecond latency, while the same feature definitions must also be usable by offline training jobs to prevent training-serving skew. Which component of a feature store architecture directly satisfies the low-latency serving requirement?

Medium
36

A media company runs an on-premises inference cluster for an image-tagging model. The model is trained on-premises and deployed into a container image that is rebuilt nightly in the company's internal registry. Security policy forbids any outbound internet access from the cluster. Which deployment approach best fits these constraints?

Medium
37

An organization wants to integrate an AI-powered summarization feature into their existing web application. The AI service will be called via API. Which factor is MOST important to consider for cost management?

Easy
38

A team is deploying a model that must comply with GDPR. Users can request deletion of their data. Which TWO practices should be implemented to support this compliance? (Select TWO.)

Medium
39

A data scientist is using PyTorch to train a custom NLP model. The training is slow on a single GPU. They want to speed up training by using multiple GPUs on a single machine. Which PyTorch feature should they use?

Medium
40

A media company stores thousands of hours of raw broadcast footage in a cloud object storage bucket. A data engineering team needs a training dataset that contains only the short clips where a goal is scored, so they must locate and extract those specific time ranges from the video files before training. Which technology should the team use to extract the required segments from the video objects?

Medium
41

A logistics company runs a route-optimization model on a fleet of delivery vehicles. Each vehicle has an NVIDIA Jetson module with limited memory, and connectivity is unavailable for hours at a time. The team wants the smallest possible runtime footprint while still executing the trained graph on the GPU. Which approach best fits these constraints?

Medium
42

A hospital's radiology department wants to run a diagnostic imaging model inside its own data center because patient images cannot leave the premises. The team needs to manage model versions, roll back a bad deployment quickly, and keep an audit trail of which model version produced each prediction. Which approach best satisfies these requirements?

Easy
43

A data science team wants to implement a feature store to serve pre-computed features for both training and inference with low latency. Which TWO tools are commonly used for building a feature store?

Medium
44

A financial services company needs to deploy an ML model for loan approval that must be explainable to regulators. The model is a gradient boosting ensemble. They need to track experiments, log model parameters, and serve the model with explanations. Which THREE tools from the MLOps ecosystem should they use?

Hard
45

A financial institution is designing an AI system to detect fraudulent transactions in real time. The system must process 10,000 transactions per second with sub-10 ms latency. The team plans to use a gradient boosting model. Which infrastructure component is most critical to meet the latency requirement?

Hard
46

A data engineering team is building a pipeline to ingest streaming user activity data, process it in real-time, and store features in a feature store for ML models. Which streaming technology is BEST suited for this real-time data ingestion and processing?

Medium
47

A healthcare analytics team is preparing to fine-tune a 7-billion-parameter open-weight language model on a single server with four NVIDIA A100 40 GB GPUs. Full fine-tuning runs out of memory, and the team wants to train on their clinical notes dataset while keeping GPU memory within the available budget. Which TWO techniques should they apply to reduce memory consumption during fine-tuning? (Choose two.)

Medium
48

A data scientist needs to store large volumes of unstructured log data for future AI model training. They also need to run SQL-based analytics on the data. Which THREE services are appropriate for this requirement? (Choose 3)

Medium
49

A hospital wants to run a natural language processing model that summarizes clinical notes. Because of patient privacy regulations, the data cannot leave the hospital's on-premises network, and there is no dedicated GPU available. Which deployment approach best fits these constraints?

Easy
50

A retail analytics team is building a retrieval-augmented generation assistant over product manuals. They need a vector index that supports fast approximate nearest neighbor search and can be updated as new manuals are published without rebuilding the entire index. Which TWO components should they use to meet these requirements? (Choose two.)

Medium
51

A financial institution is deploying a real-time anomaly detection model on a Kubernetes cluster. The model must process streaming transactions with low latency and scale horizontally during peak hours. The team wants to use a serving solution that integrates natively with Kubernetes and supports autoscaling based on request concurrency. Which solution best meets these requirements?

Hard
52

A security team needs to ensure that all data used for AI model training in the cloud is encrypted at rest and in transit. Which set of measures meets this requirement on AWS?

Medium
53

A data engineer is building a real-time feature store for an AI recommendation engine. The system must ingest millions of clickstream events per second, retain each event for 7 days, and allow the ML model to read the most recent user activity with sub-10ms latency. Which storage technology should the engineer select for the online feature serving layer?

Easy
54

A company wants to use a pre-trained model from a cloud-based AI service but must ensure that customer data is not used to improve the service. Which configuration should they choose?

Medium
55

An AI team is optimizing a convolutional neural network (CNN) for inference on a mobile device. The model has many layers and uses 32-bit floating-point weights. They need to reduce the model size and latency without significant accuracy loss. Which technique should they apply?

Hard
56

A retail analytics team is preparing a recommendation model for production. They need to serve many concurrent requests with predictable latency and also reduce the cost of running the model on GPU nodes. Which TWO practices best support these goals? (Choose two.)

Medium
57

A team uses Apache Kafka to stream real-time sensor data for ML inference. They need to process the stream, perform feature engineering, and store results in a data lake. Which tool is best suited for this streaming ML pipeline?

Medium
58

A media company wants to generate short video summaries from long recordings using a generative AI model. The model is hosted in the cloud, and the company needs to minimize cost while handling unpredictable traffic spikes. Which cloud service model is most appropriate?

Medium
59

An ML platform team is running a recommendation model on a Kubernetes cluster with GPU nodes. During peak traffic, inference pods are frequently evicted and restarted, causing latency spikes. The team wants to reduce restart frequency and keep GPU utilization high without changing the model. Which combination of Kubernetes configuration changes should they apply?

Hard
60

A hospital wants to run a diagnostic image classifier entirely inside its own data center because patient images cannot leave the premises. The IT team needs a deployment model that keeps all data and inference local while still allowing the AI team to push updated model versions. Which deployment approach fits these requirements?

Easy
61

A company is deploying a computer vision model to smartphones for offline object detection. The model was trained in PyTorch. Which format should they use for deployment on iOS devices?

Medium
62

A media company uses a large language model (LLM) to generate article summaries. They want to reduce inference costs and latency without significantly degrading summary quality. The LLM is currently served at full precision. Which optimization technique is most appropriate?

Hard
63

A data scientist wants to develop a computer vision model using transfer learning. They need a framework that provides pre-trained models and easy-to-use APIs for data augmentation and training. Which TWO frameworks are best suited for this task?

Easy
64

A hospital's AI team is deploying a real-time patient deterioration prediction model on bedside monitoring devices. The devices have limited RAM (512 MB) and no GPU, and the model must perform inference within 50 ms. The team has a trained TensorFlow model saved as a SavedModel. Which deployment approach best meets these constraints?

Medium
65

A developer wants to integrate an AI-powered text summarization API into their application. They need to authenticate securely and manage usage limits. What is the standard mechanism for authenticating with cloud-based AI services?

Easy
66

A data engineer needs to process streaming clickstream data for real-time feature engineering in an ML pipeline. Which data pipeline technology is BEST suited for this task?

Easy
67

A media company trains a video tagging model on a large dataset in the cloud. The model will run inference on-premises in a facility with intermittent network connectivity, and the operations team wants to avoid re-authoring the model for each target runtime. Which deployment artifact best meets these constraints?

Hard
68

A startup is building a conversational AI assistant that must understand and generate human-like text. The team has limited labeled data and a modest budget for compute. They want to leverage existing large language models rather than pretraining one. Which approach best meets their needs?

Easy
69

A company has a TensorFlow model trained on-premises and wants to deploy it on AWS SageMaker for scalable inference. What is the BEST way to package the model for deployment?

Medium
70

During inference, a model served via a REST API occasionally returns high latency due to cold starts. The team uses a containerized service on Kubernetes with horizontal pod autoscaling. Which solution minimizes cold start impact while controlling cost?

Hard
71

A data engineering team is designing a data pipeline to process streaming sensor data and feed it into an ML model for anomaly detection. Which THREE components are essential for this pipeline?

Medium
72

A machine learning engineer wants to track hyperparameter experiments and compare results across runs. Which TWO tools are best suited for this purpose? (Choose 2)

Medium
73

A financial institution runs a credit-scoring model that must comply with internal governance requiring that every individual prediction be traceable to the input features that drove it, and that the explanation be produced at inference time for each applicant. The model is a complex gradient-boosted ensemble. Which approach best satisfies the requirement to generate a per-prediction explanation for each applicant?

Hard
74

A research lab is training a large language model on a cluster of GPUs. They notice that training throughput decreases significantly when scaling from 8 to 16 GPUs. The model uses data parallelism with synchronous updates. Which factor is most likely causing the decreased throughput?

Hard
75

A developer wants to deploy a scikit-learn model as a REST API endpoint with minimal infrastructure management. Which cloud service is MOST appropriate?

Easy
76

A data science team uses Vertex AI for model training and deployment. They want to implement CI/CD for ML pipelines. Which THREE Google Cloud services should they integrate?

Hard
77

An AI platform team is deploying a large language model for internal document summarization. Legal requires that no prompt or document content leaves the company's virtual private cloud, and the security team wants to control the exact model weights and runtime version. The team already has GPU capacity reserved in their own VPC. Which deployment approach best satisfies these constraints?

Medium
78

A team uses Kubeflow to manage ML workflows on Kubernetes. They want to automate hyperparameter tuning for a training job. Which Kubeflow component should they use?

Medium
79

A hospital's radiology department is deploying an AI system that analyzes chest X-rays to flag potential pneumonia. Because patient data cannot leave the hospital's on-premises network, the model must run locally. The IT team wants to ensure the model's inference results can be explained to radiologists and auditors. Which approach best satisfies the explainability requirement while keeping the model on-premises?

Medium
80

A company is building a recommendation system that uses user embeddings stored in a vector database. The system must retrieve the top 10 most similar items for a given user query. Which vector database feature is MOST critical for this task?

Medium
81

A team is using a cloud AI service with a pay-per-token pricing model. They want to minimize costs while maintaining response quality. Which strategy is MOST effective?

Medium
82

An AI platform team runs inference for an image classifier on a shared GPU node. Multiple model replicas currently load the full model weights into GPU memory independently, and the node runs out of GPU memory when a third replica starts. The team wants to serve more replicas per GPU without changing model accuracy. Which approach best addresses the constraint?

Hard
83

A company is using Google Cloud Vertex AI for model training. They want to automate the retraining pipeline when new data arrives in BigQuery. Which Vertex AI feature should they use?

Medium
84

A computer vision team is preparing a model for deployment to a fleet of low-power cameras that run on battery and have limited RAM. They want to reduce model size and inference cost while keeping accuracy acceptable for detecting a small set of object classes. Which TWO techniques should they apply? (Choose two.)

Medium
85

A machine learning engineer needs to train a deep neural network on a large image dataset. Which hardware component is specifically optimized for this task due to its high parallel processing capability and is commonly used in AI training?

Easy
86

An organization must ensure that an AI model deployed on an IoT device meets stringent latency requirements. The model is currently in FP32 and runs at 200ms per inference on the device; the target is 50ms. Which technique will provide the greatest latency reduction with the least accuracy loss?

Hard
87

An AI research group trains a large language model across a cluster of GPU nodes. They observe that training throughput drops sharply whenever gradient synchronization occurs, and profiling shows GPUs idle while waiting for parameter updates to be exchanged. The model must remain mathematically identical to single-node training. Which change should the team make?

Hard
88

A retail chain is deploying an AI-powered demand forecasting system across 500 stores. The system ingests daily sales, weather, and promotion data, and must produce forecasts that update as new data arrives. The MLOps team needs to ensure the deployed model remains accurate over time as consumer behavior shifts. Which TWO practices should they implement? (Choose two.)

Medium
89

A company wants to build an AI pipeline that processes streaming data from IoT sensors, performs feature engineering, trains a model incrementally, and deploys the updated model. Which data pipeline technology is BEST suited for the streaming ingestion step?

Medium
90

A machine learning engineer is designing a pipeline to train a computer vision model using PyTorch on a large dataset stored in an S3 data lake. They need to preprocess images (resize, normalize) and stream them efficiently to GPUs. Which THREE components are essential in this pipeline? (Select THREE.)

Hard
91

An AI team wants to version control datasets, track experiments, and log model parameters across multiple projects. Which MLOps platform is specifically designed for experiment tracking and model management?

Easy
92

Which of the following is a key advantage of using ONNX (Open Neural Network Exchange) format for model deployment?

Easy
93

A financial institution is deploying an AI model for credit scoring. The model must be explainable to regulators, and the team needs to understand which features contribute most to individual predictions. Which TWO techniques should they use? (Choose two.)

Medium
94

A media company wants to automatically generate concise summaries of lengthy earnings-call transcripts. The transcripts average 45 minutes of speech and contain domain-specific financial terminology. The team needs a solution that captures long-range dependencies and produces fluent, abstractive summaries without training a model from scratch. Which approach is most appropriate?

Hard
95

Which open-source framework is commonly used for building, training, and deploying machine learning models and provides high-level APIs like Keras?

Easy
96

A team is deploying a machine learning model on a Kubernetes cluster. They need to ensure low-latency inference and efficient resource utilization. Which approach should they use to dynamically scale inference pods based on request volume?

Hard
97

A company needs to store large volumes of unstructured data (PDFs, images, logs) for future AI model training. The data must be easily accessible by data scientists using Spark and must support cost-effective storage. Which data infrastructure is MOST appropriate?

Medium
98

An AI platform team is building a retrieval-augmented generation service over an internal knowledge base of roughly 40 million technical documents. Queries must return semantically relevant passages in under 50 ms at the vector search layer. The team wants approximate nearest neighbor search that supports metadata filtering on fields such as product line and document date, and they want to avoid a separate relational database for those filters. Which vector index type best matches these requirements?

Hard
99

Which hardware accelerator is specifically designed by Google for training and inference of machine learning models, particularly their TensorFlow framework?

Easy
100

An ML team uses Kubeflow to orchestrate a pipeline that includes data preprocessing, model training, and evaluation. The pipeline runs on a Kubernetes cluster. After a cluster upgrade, the pipeline fails at the training step with an 'OOMKilled' error. What is the MOST likely cause?

Hard
101

A data engineer is building a pipeline to process streaming clickstream data and feed it into a real-time ML feature store. Which tool is BEST suited for the streaming ingestion?

Medium
102

A company is building a secure AI system that must comply with GDPR. They want to allow users to request deletion of their personal data from training sets and model outputs. Which THREE techniques should they implement?

Hard
103

A team is building a retrieval-augmented generation (RAG) pipeline. They need to store embeddings of company documents and perform fast similarity searches. Which data store is BEST suited for this task?

Medium
104

A developer is building a mobile app that uses a pre-trained image classification model on-device. Which framework should they use to run the model on iOS devices?

Easy
105

A logistics company runs a route-optimization model on a Kubernetes cluster. During peak hours the inference pods are frequently evicted because the nodes run out of memory, even though average GPU utilization stays below 40 percent. The team wants to reduce evictions without changing the model or adding nodes. Which action best addresses the root cause?

Hard
106

A startup is training a recommendation model on a single workstation with one GPU. The dataset has grown to 2 TB, and training now takes several days. The team wants to reduce training time by adding more GPUs to the same workstation. Which technology should they use to enable efficient multi-GPU training with minimal code changes?

Easy
107

A computer vision team trains a convolutional neural network for manufacturing defect detection on a workstation with an NVIDIA RTX A6000 GPU. They want to reduce training time by increasing throughput without changing model architecture or batch size. Which action should they take?

Medium
108

Which AI accelerator is specifically designed by Google to accelerate the training and inference of large neural networks, especially in their cloud environment?

Easy
109

An organisation is deploying a fine-tuned LLM for internal use. They need to ensure the API endpoint is secure and cost-effective. Which TWO measures should they implement? (Choose 2)

Hard
110

An MLOps engineer is deploying a scikit-learn random forest model to a Kubernetes cluster for a low-traffic internal API. The team wants to avoid maintaining a custom Flask wrapper and prefers a standard serving solution that supports REST and gRPC. Which serving component should they choose?

Hard
111

A platform team is preparing a feature store for a recommendation system. They need point-in-time correct feature retrieval so that training datasets do not leak future information, and they need the same features served online with low latency. Which architecture best satisfies both requirements?

Hard
112

A data science team is deploying a real-time fraud detection model on edge devices in retail stores. The model must infer under 10ms and fit within 50MB memory. Which combination of techniques should the team apply?

Hard
113

A financial institution is deploying an AI model that predicts loan default risk. The model is trained on historical data that includes sensitive attributes like zip code and marital status. The compliance team is concerned about disparate impact. Which technique should be applied during model training to mitigate bias while maintaining predictive performance?

Medium
114

A startup is developing an AI chatbot and wants to use a pre-trained language model to generate responses. They need to integrate the model into their application with minimal latency and cost. Which approach should they take?

Easy

Frequently asked questions

What does the AI Infrastructure and Technologies domain cover on the AI0-001 exam?
Be able to select accelerators, deployment architectures, streaming tools, and model packaging formats from stated constraints. The single most important skill is mapping the constraint — latency, compliance, cloud vendor, or data velocity — to the correct infrastructure choice, not the most familiar one.
How many questions are in this domain?
This page lists all 114 AI Infrastructure and Technologies questions in the AI0-001 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only AI Infrastructure and Technologies questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
comptia-ai COMPTIA-AI aio ai infrastructure Practice Questions