Courseiva

AI-900 · domain

Describe fundamental principles of machine learning on Azure

Use this page to practise Describe fundamental principles of machine learning on Azure questions for this certification. Focus on how the exam tests describe fundamental principles of machine learning on azure in scenario format — understanding the why behind each answer builds more durable knowledge than memorising options.

200 questions47 easy125 medium28 hard

Focused practice

Practice Describe fundamental principles of machine learning on Azure questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Describe fundamental principles of machine learning on Azure

Describe fundamental principles of machine learning on Azure questions on this certification test your ability to deploy and manage describe fundamental principles of machine learning on azure concepts in scenario-based situations.

Core Describe fundamental principles of machine learning on Azure concepts and how they apply in real-world cloud scenarios.

How to deploy describe fundamental principles of machine learning on azure correctly and verify the outcome.

Troubleshooting describe fundamental principles of machine learning on azure issues by interpreting error output and system state.

Cloud best practices and Describe fundamental principles of machine learning on Azure design trade-offs tested by this certification.

Watch out for

Common Describe fundamental principles of machine learning on Azure exam traps

  • Selecting the most expensive service when a simpler managed option meets the requirement.
  • Forgetting that cloud resources must be explicitly secured — defaults are rarely secure.
  • Choosing a global service fix when the issue is region-specific.
  • Overlooking cost implications of cross-region data transfer in architecture questions.

Question index

All Describe fundamental principles of machine learning on Azure questions (200)

Click any question to see the full explanation, or start a practice session above.

1

What is the difference between a binary classification model and a multi-class classification model?

Medium
2

What is 'Bayesian optimisation' in hyperparameter tuning?

Medium
3

A data scientist wants to train a model that predicts whether a customer will respond to a marketing offer (yes or no). The dataset includes features such as age, income, past purchase history, and the labeled outcome (responded or not responded) for previous customers. Which type of machine learning is this?

Easy
4

What does it mean for an ML model to 'generalize'?

Medium
5

A data science team trains a regression model to predict house prices. They evaluate the model using Mean Absolute Error (MAE). After deployment, they notice that the model occasionally produces large errors (e.g., underpredicting a luxury home by $500,000) while most predictions are within $20,000. The business is more concerned about the impact of these large errors than the average small error. Which additional metric should the team use to better capture the penalty for large errors?

Hard
6

A retail company wants to analyze customer purchase histories to identify natural groups of customers with similar buying patterns. They do not have predefined categories. Which type of machine learning should they use?

Medium
7

What is 'model lineage' in Azure Machine Learning?

Medium
8

What is hyperparameter tuning in machine learning?

Medium
9

A retail company wants to predict the exact number of units of a product that will be sold next month. They have historical sales data and information about promotions and holidays. The target variable is the number of units sold, which is a continuous value. Which type of machine learning task should they perform?

Easy
10

A data scientist trains a regression model to predict housing prices. The model uses polynomial features up to degree 5. It achieves an R-squared of 0.95 on the training set but only 0.60 on the test set. Which problem is the model most likely experiencing?

Hard
11

A data scientist trains a deep neural network on a small dataset. The model achieves 100% accuracy on the training data but only 60% accuracy on a validation set. Which technique is most appropriate to address this issue?

Medium
12

What is 'training data' vs 'test data' in machine learning?

Easy
13

A data scientist trains a model to predict house prices. The model achieves 99% accuracy on the training data but only 80% accuracy on new test data. Which technique is most likely to help improve the model's generalization?

Medium
14

What is 'model versioning' and why is it essential in MLOps?

Medium
15

What is 'gradient boosting' and how does it differ from random forests?

Hard
16

A data scientist trains a machine learning model on historical sales data to predict future sales volume. The model achieves 99% accuracy on the training dataset but only 75% accuracy on a separate test dataset. What is the most likely issue with this model?

Medium
17

A robotics team is training a robot to navigate a maze. The robot receives a positive reward (+10) when it reaches the exit and a negative reward (-1) every time it bumps into a wall. The robot learns to maximize its cumulative reward over multiple trials. Which type of machine learning is being used?

Medium
18

What is 'ensemble learning' in machine learning?

Medium
19

A data scientist trains a classification model to predict whether an email is 'phishing' or 'legitimate'. The model achieves 99% accuracy on the training data but only 68% accuracy on the test data. Which action is most likely to help improve the model's generalization performance?

Medium
20

What is 'Azure Machine Learning workspace' and what does it contain?

Easy
21

What does 'deep learning' refer to in machine learning?

Easy
22

A data scientist has a small dataset with only 200 labeled samples. They want to get a reliable estimate of model performance without using a separate validation set that would reduce the training data. Which technique should the data scientist use in Azure Machine Learning to obtain this reliable estimate?

Hard
23

What is 'cross-validation' and when should it be used in machine learning?

Medium
24

What is the purpose of splitting data into training, validation, and test sets in machine learning?

Medium
25

What is 'ensemble learning' in machine learning and why does it improve performance?

Medium
26

A data scientist trains a binary classification model to detect fraudulent transactions. The dataset contains 99% legitimate transactions (negative class) and 1% fraudulent transactions (positive class). The model predicts 'legitimate' for every transaction in the test set and achieves 99% accuracy. Which metric would best reveal that the model is failing to identify any fraudulent transactions?

Medium
27

Drag and drop the steps to train a custom vision model in Azure Custom Vision into the correct order.

Medium
28

A data scientist trains a linear regression model to predict house prices. The model's training error is very high, and its test error is nearly as high. Which term best describes this situation?

Medium
29

What is 'differential privacy' and how is it relevant to AI model training?

Hard
30

A data scientist is training a binary classification model to detect fraudulent transactions. The dataset contains only 1% fraudulent transactions. The model achieves 99% accuracy on the test set, but when deployed, it fails to detect most actual fraud cases. Which metric would best reveal this issue?

Hard
31

What is 'time series forecasting' and what Azure ML tools support it?

Medium
32

A data scientist evaluates a regression model that predicts house prices. On the test set, the Mean Absolute Error (MAE) is $8,000 and the Root Mean Squared Error (RMSE) is $25,000. What does the large difference between MAE and RMSE indicate about the model's errors?

Hard
33

A data scientist is building a classification model to predict customer churn. The dataset has only 5% churn cases. The model achieves 95% accuracy on the test set, but upon investigation, the data scientist finds the model predicts 'not churn' for nearly every customer. Which metric should the data scientist primarily use to evaluate the model's performance on this imbalanced dataset?

Medium
34

A data scientist trains a multiclass classification model to identify different species of flowers (Iris setosa, Iris virginica, Iris versicolor). The overall accuracy is 94%, but the accuracy for the Iris virginica class is only 60%. Which additional metric should the data scientist examine to better understand the model's performance on the minority class?

Medium
35

Match each Azure AI service to its pricing model.

Medium
36

What is the 'bias-variance tradeoff' in machine learning?

Hard
37

A retail company wants to automatically group customers into segments based on their purchasing history, age, and location without using any predefined labels. The goal is to identify distinct customer profiles for targeted marketing campaigns. Which type of machine learning approach should they use?

Easy
38

A data scientist trains a regression model to predict house prices. The model achieves very low error on the training data but significantly higher error on a held-out test set. Which problem does this scenario best describe?

Medium
39

What is 'hyperparameter tuning' in Azure Machine Learning?

Medium
40

A robotics company is training a drone to fly autonomously through an obstacle course. The drone receives positive rewards for staying on course and avoiding obstacles, and negative rewards for collisions. The system learns by trial and error to maximize its cumulative reward. Which type of machine learning is being used?

Easy
41

A city's traffic department wants to predict the number of cars that will cross a particular bridge each day to plan maintenance schedules. The output of the model should be a numerical value representing the estimated traffic count. Which type of machine learning task is this?

Medium
42

What is 'Azure Machine Learning pipelines' and why are they used?

Medium
43

A data scientist has a dataset containing information about houses: size (sq ft), number of bedrooms, location, and the actual sale price. The goal is to train a model that predicts the price of a new house based on these features. Which type of machine learning task is this?

Medium
44

What is 'model registry' in Azure Machine Learning?

Medium
45

Which Azure service provides a no-code/low-code drag-and-drop interface for building machine learning pipelines?

Medium
46

A data scientist trains a classification model to predict whether an email is spam or not. The model achieves 98% accuracy on the test set, but upon inspection, it classifies all emails as 'not spam' because the dataset has 95% non-spam emails. What is the most likely issue?

Medium
47

Drag and drop the steps to analyze an image with Azure Computer Vision into the correct order.

Medium
48

A data scientist trains a classification model on a dataset of 10,000 labeled emails to distinguish spam from non-spam. The model achieves 99% accuracy on the training data but only 70% accuracy on a held-out test set. Which term best describes this situation?

Medium
49

A real estate company has a dataset containing square footage, number of bedrooms, and location for 10,000 houses, along with their sale prices. They want to train a model that predicts the sale price of a new house based on these features. Which type of machine learning should they use?

Easy
50

What is the role of a label (also called target or ground truth) in supervised machine learning?

Easy
51

What is 'dimensionality reduction' and why is it useful in machine learning?

Easy
52

What is the bias-variance tradeoff in machine learning?

Hard
53

A real estate company trains a model to predict house prices. They evaluate it on a test set of 100 houses. The model predictions have a mean absolute error (MAE) of $5,000 and a root mean squared error (RMSE) of $20,000. What does the large difference between MAE and RMSE indicate about the model's errors?

Hard
54

A data scientist is building a classification model to detect fraudulent transactions. The dataset has 1,000,000 legitimate transactions and only 1,000 fraudulent ones. The model achieves 99.9% accuracy on the test set, but it fails to catch most fraudulent cases. Which metric should the data scientist prioritize to better evaluate the model's performance on this imbalanced dataset?

Hard
55

A company builds a machine learning model to predict whether a customer will purchase a product. They use a training dataset with 50% purchasers and 50% non-purchasers. The model achieves 90% accuracy on the test set. However, when deployed, the model performs poorly because the actual customer base has only 5% purchasers. What is the most likely cause of this poor performance?

Easy
56

What is the purpose of a confusion matrix in evaluating a classification model?

Medium
57

A data scientist trains a regression model to predict house prices. The model has a mean absolute error (MAE) of $5,000 on the test set. Which statement best interprets this metric?

Medium
58

A data scientist trains a decision tree model to predict customer churn. The model achieves 99% accuracy on the training data but only 80% on the test data. Which concept best explains this performance difference?

Medium
59

What is 'active learning' in Azure Machine Learning data labelling?

Medium
60

A data scientist trains a binary classification model to detect spam emails. The dataset contains 95% legitimate emails (negative class) and 5% spam (positive class). The model predicts all emails as legitimate. The accuracy is 95%, but the model is useless. Which metric would best indicate the model's failure?

Hard
61

What is the purpose of Azure Machine Learning's automated ML (AutoML) feature?

Medium
62

A data scientist uses Azure Machine Learning to train a model that predicts the electricity consumption (in kilowatt-hours) of a building based on features like building age, square footage, and number of occupants. The data scientist wants to evaluate how accurately the model's predictions match the actual consumption values. Which evaluation metric is most appropriate for this regression task?

Easy
63

What is the Azure Machine Learning model registry?

Medium
64

What is 'regularization' in machine learning and why is it used?

Medium
65

What is 'model interpretability' and which Azure tool helps with it?

Medium
66

What is the difference between 'precision' and 'recall' as model evaluation metrics?

Hard
67

A data scientist has trained a binary classification model to detect fraudulent credit card transactions. The dataset contains 99.9% legitimate transactions and only 0.1% fraudulent ones. The model predicts all transactions as legitimate, achieving 99.9% accuracy on the test set. However, the business requires the model to actually catch as many fraudulent transactions as possible. Which metric would best reveal the model's failure to identify fraud?

Medium
68

A retail company has historical data about customers, including age, purchase history, and whether they have churned (yes/no). They want to train a model that predicts if a new customer will churn. Which type of machine learning should they use?

Easy
69

What is an ML pipeline in Azure Machine Learning?

Medium
70

A data scientist is training a model to classify customer reviews as positive, negative, or neutral. The dataset contains 10,000 reviews, but only 500 of them are negative. The data scientist wants to ensure the model performs well on the minority class (negative reviews). Which technique should the data scientist consider to address the class imbalance?

Medium
71

What is 'Azure Machine Learning environments' and why are they important for reproducibility?

Medium
72

A data scientist has a dataset containing thousands of labeled images of cats and dogs. The data scientist wants to train a model that can automatically classify new unlabeled images as either 'cat' or 'dog'. Which type of machine learning should the data scientist use?

Easy
73

What is the 'mean absolute error' (MAE) metric used to evaluate in machine learning?

Easy
74

A data scientist trains a binary classification model to predict loan defaults. The dataset contains 98% non-default cases and only 2% default cases. The model predicts 'non-default' for every instance, achieving 98% accuracy on the test set. Which metric would best reveal that the model fails to identify any actual defaults?

Hard
75

Match each Azure AI service to its associated API or SDK.

Medium
76

An e-commerce company has a dataset of customer purchase histories with no predefined categories. The data analyst wants to identify natural groupings of customers based on their purchasing behavior to target marketing campaigns. Which type of machine learning should the analyst use?

Easy
77

A bank uses a machine learning model to predict credit card fraud. The model's output is a probability score. The business wants to minimize the number of false positives (legitimate transactions incorrectly flagged as fraud) because these cause customer dissatisfaction. At the same time, they must also catch most fraudulent transactions. Which metric should the bank optimize to balance these two goals?

Hard
78

A data scientist has a dataset with 100 features and 10,000 samples. They want to reduce the number of features while retaining as much variance as possible, to improve model training speed and reduce overfitting. Which technique should they use?

Medium
79

What is 'ONNX' and why is it relevant to Azure AI?

Medium
80

A data scientist trains a regression model to predict house prices using features like bedrooms, square footage, and location. The model achieves a low error on the training data but performs significantly worse when used to predict prices in a new city with different property characteristics. Which concept best explains this poor performance?

Medium
81

What is 'semi-supervised learning' and when is it useful?

Medium
82

What is 'overfitting' in machine learning and how does Azure ML help prevent it?

Medium
83

What is a training job in Azure Machine Learning?

Medium
84

What is 'model deployment' in Azure Machine Learning?

Easy
85

What is a neural network?

Medium
86

What is 'ROC-AUC' and when is it a better metric than accuracy for classification?

Medium
87

A hospital deploys a machine learning model to screen patients for a rare disease. Only 0.1% of patients actually have the disease. The model correctly identifies most positive cases but also flags many healthy patients as potentially having the disease. The hospital wants to minimize the number of healthy patients who are incorrectly told they might have the disease. Which metric should the model optimize?

Medium
88

A data scientist is preparing a dataset to train a model that predicts customer churn. The dataset includes a column 'CustomerID' which is a unique identifier for each customer. Should the data scientist include the 'CustomerID' column as a feature in the training data?

Easy
89

A data scientist trains a machine learning model to predict house prices based on features like square footage, number of bedrooms, and location. The model achieves a very low error on the training data but performs poorly on a held-out test set. Which term best describes this situation?

Medium
90

Which metric is MOST appropriate for evaluating a regression model's performance?

Medium
91

A data scientist is building a binary classification model to predict fraudulent credit card transactions. The dataset is highly imbalanced: only 1% of transactions are fraudulent. The cost of a false negative is very high because missing a fraudulent transaction can lead to significant financial loss. Which evaluation metric should the data scientist prioritize to minimize false negatives?

Medium
92

A data scientist wants to train a machine learning model to predict the exact market price of a house based on features such as square footage, number of bedrooms, and location. Which type of machine learning task should be used?

Medium
93

A data scientist trains a binary classification model to detect fraudulent credit card transactions. The dataset contains 99.5% legitimate transactions and 0.5% fraudulent transactions. The model predicts every transaction as legitimate and achieves 99.5% accuracy on the test set. Which metric would best reveal that the model is failing to identify any fraudulent transactions?

Medium
94

A data scientist is developing a classification model to detect fraudulent transactions. The dataset is split into training and test sets. The data scientist repeatedly tunes the model's hyperparameters and evaluates performance on the test set until the test accuracy reaches 95%. However, when the model is deployed on new, unseen data, its accuracy drops to 70%. Which concept best explains this performance degradation?

Medium
95

What is an endpoint in Azure Machine Learning?

Easy
96

What does 'model accuracy' measure in machine learning classification?

Medium
97

A medical research team trains a model to detect a rare disease from lab results. The disease occurs in only 1% of patients. The model predicts 'no disease' for every patient and achieves 99% accuracy. Which metric best reveals that the model is failing to identify actual disease cases?

Medium
98

A data scientist trains a model to predict house prices using features like number of bedrooms, square footage, and location. The model achieves a mean absolute error (MAE) of $5,000 on the training data but $25,000 on the test data. Which problem is the model most likely experiencing?

Easy
99

A data scientist trains a model to predict the exact number of cars that will cross a bridge each day for maintenance planning. The model uses historical traffic data as input. Which type of machine learning task is this?

Medium
100

What is 'model evaluation' and what metrics are used for different ML task types?

Medium
101

What is 'regularisation' in machine learning and what problem does it solve?

Medium
102

A data scientist is evaluating a binary classification model that predicts whether a transaction is fraudulent. The test set contains 1,000 transactions: 990 legitimate and 10 fraudulent. The model's predictions are shown in the confusion matrix below. Confusion matrix: Predicted Legitimate Predicted Fraudulent Actual Legitimate 942 48 Actual Fraudulent 2 8 Which metric should the data scientist prioritize if the business goal is to minimize the number of fraudulent transactions that are missed (false negatives)?

Hard
103

A data scientist has a dataset containing customer transaction records with features such as age, income, and purchase history, but no labels. The goal is to identify natural groupings of customers for a targeted marketing campaign. Which type of machine learning should be used?

Medium
104

What is 'Azure Machine Learning Responsible AI dashboard's error analysis'?

Easy
105

What is 'imbalanced classification' handling using 'SMOTE'?

Medium
106

What is a 'compute instance' in Azure Machine Learning?

Medium
107

A data scientist is training a regression model to predict house prices using features like square footage, number of bedrooms, and location. After evaluating the model on a test set, the data scientist wants to select a metric that measures the average magnitude of prediction errors in the same units as the target variable (price). Which evaluation metric should the data scientist use?

Medium
108

What is 'online learning' (incremental learning) in machine learning?

Medium
109

A data scientist is training a model to predict whether a patient has a rare disease (1% prevalence). The model predicts 'no disease' for all patients and achieves 99% accuracy, but fails to identify any actual cases. Which metric would best reveal this failure?

Medium
110

What is 'Azure Machine Learning compute' and what types are available?

Easy
111

A data scientist is training a regression model to predict house prices. The data scientist wants to evaluate the model using a metric that penalizes large prediction errors significantly more than small errors. Which evaluation metric should the data scientist choose?

Medium
112

Which type of machine learning uses labeled training data where the correct output is provided for each input?

Medium
113

A data scientist is building a model to predict the exact temperature in degrees Celsius based on humidity and atmospheric pressure. The model will output a single numeric value for each input. Which type of machine learning task is this?

Easy
114

What is 'transfer learning' and how is it different from training from scratch?

Medium
115

What is 'k-fold cross-validation' specifically and how is k=10 different from k=5?

Medium
116

A data scientist trains a binary classification model to detect fraudulent transactions. The dataset contains only 2% fraudulent transactions. The model achieves 98% overall accuracy, but it fails to detect any fraudulent transactions, classifying all transactions as legitimate. Which metric would most clearly reveal this failure?

Hard
117

A data scientist is training a binary classification model to detect rare equipment failures from sensor data. The dataset contains 99.5% normal operation readings and only 0.5% failure readings. The model currently predicts all readings as 'normal' and achieves 99.5% accuracy on the test set. The business requires the model to identify at least 80% of actual failures. Which data-level technique should the data scientist use to most directly address the class imbalance?

Hard
118

What is feature engineering in machine learning?

Easy
119

What is the F1 score in machine learning evaluation?

Medium
120

What is overfitting in machine learning?

Medium
121

A data scientist wants to group customers into segments based on purchasing behavior without using any labeled examples. Which type of machine learning is this?

Easy
122

A data scientist trains a model on historical data and achieves high accuracy on both the training set and a held-out test set. However, when the model is deployed in production, it performs poorly on new, unseen data. Which issue is most likely the cause?

Medium
123

A hospital has a dataset with historical patient records, each labeled as either 'readmitted within 30 days' or 'not readmitted'. The hospital wants to train a model to predict which current patients are likely to be readmitted. Which type of machine learning task is this?

Medium
124

What is 'Azure Machine Learning's Responsible AI dashboard' and what does it include?

Medium
125

What is 'automated machine learning' (AutoML) in Azure Machine Learning?

Easy
126

What is 'causal inference' and how does it differ from correlation-based machine learning?

Hard
127

What is 'batch inference' vs 'real-time inference' in Azure Machine Learning?

Medium
128

What is 'Azure Machine Learning notebooks' and who typically uses them?

Easy
129

What is data drift in the context of deployed machine learning models?

Medium
130

A media company wants to automatically organize a large collection of news articles into several topic-based categories (e.g., politics, sports, technology) without using any predefined labels. They plan to use Azure Machine Learning. Which type of machine learning task should they use?

Medium
131

A data scientist trains a regression model on a dataset with 100 features and 10,000 samples. The model achieves a low training error but a much higher error on a held-out test set. Which approach is most likely to improve the model's generalization performance?

Medium
132

A data scientist trains a machine learning model on a dataset of housing prices. The model achieves 98% accuracy on the training data but only 72% accuracy on a separate test set. What is the most likely problem with this model?

Medium
133

A data scientist trains a binary classification model to distinguish between images of cats and dogs. On the test set, the model achieves 98% accuracy, but a deeper inspection reveals that the test set contains 95% cats and 5% dogs, and the model predicts 'cat' for every single image. Which metric should the data scientist prioritize to get a more realistic evaluation of the model's performance on this imbalanced dataset?

Easy
134

A data scientist has a dataset containing images of handwritten digits (0-9) where each image is labeled with the correct digit. The goal is to train a model that can predict the digit from a new image. Which type of machine learning approach should be used?

Medium
135

A data scientist trains a regression model to predict house prices. The model performs poorly on both the training data and the test data, showing high error in both sets. Which concept best describes this situation?

Medium
136

What is 'Azure Machine Learning datasets' and why are they important?

Medium
137

A data scientist is training a regression model to predict house prices in Azure Machine Learning. The model uses features like square footage, number of bedrooms, and location (zip code). The data scientist notices that the model has a very low error on the training data but a high error on the test data. Which technique should the data scientist apply during model training to reduce overfitting by penalizing large coefficients?

Medium
138

What is 'stochastic gradient descent' (SGD) and how does it work?

Medium
139

A data scientist trains a regression model to predict daily electricity consumption (in kWh) for a commercial building. The business team needs a metric that heavily penalizes large prediction errors (outliers) more than small errors. Which metric should the data scientist report to best meet this requirement?

Medium
140

A data scientist is training a classification model on a dataset with 100 features and only 500 labeled samples. The model achieves 99% accuracy on the training data but only 68% accuracy on a held-out test set, indicating overfitting. Which technique is most appropriate to directly address this problem?

Medium
141

What is 'label imbalance' in a classification dataset and how does it affect model training?

Medium
142

A data scientist is using Azure Automated Machine Learning to build a binary classification model for a highly imbalanced dataset (95% negative, 5% positive). The data scientist wants AutoML to select the best model based on a metric that is robust to class imbalance. Which primary metric should the data scientist configure in the AutoML settings?

Hard
143

What is model monitoring in Azure Machine Learning and why is it important?

Medium
144

What is 'model monitoring' in Azure Machine Learning after deployment?

Medium
145

What is 'neural architecture search' (NAS) and how does it relate to AutoML?

Hard
146

A bike-sharing company wants to predict the number of rentals per hour. Their model's predictions are usually close but occasionally have large errors due to unexpected events like sudden rain. They want a metric that heavily penalizes these large errors to ensure the model is not overly confident. Which evaluation metric should they primarily use?

Medium
147

What is 'Azure ML's experiment tracking' and why do data scientists use it?

Easy
148

What is machine learning?

Easy
149

A manufacturer trains a model to detect defective parts on an assembly line. Only 2% of parts are defective. The model predicts 'non-defective' for all parts and achieves 98% accuracy. Which metric best reveals the model's inability to identify defective parts?

Medium
150

What is 'curriculum learning' and how does it relate to training stability?

Hard
151

What is 'data augmentation' and how does it help with limited training data?

Medium
152

What is 'clustering' in unsupervised machine learning?

Easy
153

What is 'regression' in machine learning and when is it used?

Easy
154

A manufacturing team wants to predict product defects based on sensor readings from the production line. They have 10,000 historical samples, each labeled as 'defective' or 'non-defective'. Which type of machine learning should they use in Azure Machine Learning?

Medium
155

What is 'data preprocessing' and why is it important for machine learning?

Easy
156

What is recall (sensitivity) in the context of binary classification model evaluation?

Medium
157

What is 'model explainability' using SHAP values in Azure Machine Learning?

Medium
158

What is 'model compression' and what techniques does it include?

Medium
159

What is Azure Machine Learning?

Easy
160

A data scientist is training a logistic regression model to predict customer churn using a small dataset with 500 records and 200 features. The model achieves 97% accuracy on the training set but only 65% on a held-out test set, indicating severe overfitting. The data scientist wants to reduce overfitting by automatically eliminating irrelevant features. Which technique should the data scientist apply?

Hard
161

A retail company wants to automatically group its customers into distinct segments based on their purchasing patterns, without having pre-defined categories. The goal is to discover natural groupings in the customer data to tailor marketing campaigns. Which type of machine learning task should the company use?

Easy
162

What is the Azure Machine Learning workspace?

Easy
163

What is 'confusion matrix' and what does it tell you about a classification model?

Medium
164

A data scientist trains a regression model to predict the selling price of houses. After evaluating on a test set, the data scientist wants a metric that measures the average absolute error between predicted and actual prices, expressed in the same units (dollars) as the target variable. Which evaluation metric should the data scientist use?

Medium
165

What is 'feature engineering' and why does it matter for machine learning models?

Medium
166

A data scientist is building a machine learning model to predict the number of daily bike rentals in a city based on weather data and day of the week. The target variable is a continuous integer. Which type of machine learning task is this?

Medium
167

A data science team trains several machine learning models for a regression task. They observe that Model A has low training error and low test error. Model B has low training error but high test error. Model C has high training error and high test error. Which model would most likely benefit from an ensemble technique that averages the predictions of multiple models?

Hard
168

A retail company has a dataset of customer transaction records with no predefined categories. They want to identify natural groupings of customers based on their purchasing behavior to create targeted marketing campaigns. Which type of machine learning should they use in Azure Machine Learning?

Medium
169

A retail company wants to predict which customers are likely to stop using their service. They have a dataset with many customer attributes including age, income, purchase history, website activity, and support interactions. They suspect some features are redundant. Which technique should they use to reduce the number of features while preserving as much information as possible?

Medium
170

What is Azure Machine Learning's 'responsible AI dashboard'?

Easy
171

What is the role of a validation dataset in machine learning?

Medium
172

A data scientist trains a binary classification model to detect fraudulent transactions. The dataset contains only 1% fraudulent cases. The model predicts 'not fraudulent' for all transactions and achieves 99% accuracy. Which metric would best reveal the model's poor performance on fraud detection?

Medium
173

A data scientist has trained a binary classification model to predict whether an email is spam (positive) or not spam (negative). On a test set, the model correctly identifies 90 out of 100 actual spam emails and 80 out of 100 actual non-spam emails. Which metric shows the proportion of actual spam emails that the model correctly predicted?

Medium
174

What is a training dataset in machine learning?

Easy
175

A data scientist is training a regression model to predict energy consumption. The dataset includes features like temperature, humidity, time of day, and day of week. After training, the model performs well on the training set but poorly on new data. Which approach would most likely help reduce this problem?

Medium
176

What is 'feature importance' in Azure Machine Learning and how is it used?

Medium
177

What is a feature in the context of machine learning?

Medium
178

What does Azure Machine Learning's 'compute cluster' provide?

Easy
179

A retail company wants to segment its customers into different groups based on purchasing behavior, without using predefined categories. Which type of machine learning task should they use?

Medium
180

A data scientist is building a machine learning model to predict whether a credit card transaction is fraudulent or legitimate. The dataset contains 100,000 historical transactions, each labeled as 'fraudulent' or 'legitimate'. Which type of machine learning task should the data scientist use in Azure Machine Learning?

Medium
181

What type of machine learning model is used for time series forecasting?

Easy
182

What is the purpose of a 'validation dataset' in machine learning?

Easy
183

What is 'Azure Machine Learning's job submission' and what types of training jobs are supported?

Medium
184

An online retailer wants to build a recommendation system that learns from user interactions. The system suggests a product, and if the user clicks it, it receives a positive reward; if ignored, a negative reward. Over time, the system learns to make better suggestions. Which type of machine learning best describes this approach?

Hard
185

A data scientist trains a regression model to predict house prices using features like bedrooms, square footage, and location. The model achieves an R-squared of 0.95 on the test set. However, when deployed to predict prices in a new city with different property characteristics, the predictions are very inaccurate. Which concept best explains this poor performance?

Hard
186

What is the purpose of Azure Machine Learning's dataset versioning?

Medium
187

What is the difference between 'training' and 'inference' in machine learning?

Easy
188

What is a confusion matrix's 'false positive' in medical screening?

Medium
189

What is 'federated learning' and when is it used for privacy-preserving AI?

Hard
190

What is cross-validation in machine learning?

Medium
191

A data scientist trains a regression model to predict energy consumption for a smart building. The model achieves very low error on the training data but performs significantly worse on a held-out validation set. Which technique would most directly address this problem?

Hard
192

A data scientist trains a machine learning model to predict housing prices. On the training data, the model achieves an R-squared value of 0.99, but on a separate validation dataset it achieves an R-squared of only 0.65. What is the most likely issue with this model?

Easy
193

A data scientist trains a binary classification model to detect a rare disease. The dataset contains 99% negative cases and only 1% positive cases. The model predicts all cases as negative, achieving an accuracy of 99% on the test set. However, the business requires the model to identify as many positive cases as possible. Which metric should the data scientist examine to best reveal that the model is failing to identify any positive cases?

Medium
194

What is 'Azure Machine Learning designer' and who is it designed for?

Easy
195

What is reinforcement learning?

Medium
196

What is AutoML in Azure Machine Learning and what does it automate?

Easy
197

What is the purpose of a test dataset in machine learning model development?

Easy
198

A data scientist trains a classification model to distinguish between images of cats and dogs. The model achieves 99% accuracy on the training set but only 75% accuracy on a validation set. Which concept best describes this situation?

Easy
199

What is precision in the context of binary classification model evaluation?

Medium
200

A data scientist is training a model to predict whether a customer will purchase a product (Yes/No). The dataset contains 90% 'No' and 10% 'Yes'. After training, the model achieves 90% accuracy. Which evaluation metric would be more informative to assess the model's performance on the minority class?

Medium

Frequently asked questions

What does the Describe fundamental principles of machine learning on Azure domain cover on the AI-900 exam?
Describe fundamental principles of machine learning on Azure questions on this certification test your ability to deploy and manage describe fundamental principles of machine learning on azure concepts in scenario-based situations.
How many questions are in this domain?
This page lists all 200 Describe fundamental principles of machine learning on Azure questions in the AI-900 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Describe fundamental principles of machine learning on Azure questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
ai-900 AI-900 describe fundamental principles of machine learning on azure Practice Questions