Courseiva

AI-900 · domain

Describe fundamental principles of machine learning on Azure

This domain covers core machine learning concepts on Azure: regression, classification, clustering, and the Azure Machine Learning workflow. The exam tests whether you can choose the right model type for a scenario, identify Azure Machine Learning components like compute, datasets, and pipelines, and distinguish automated ML from designer. Expect scenario-based questions rather than deep coding.

199 questions47 easy126 medium26 hard

Focused practice

Practice Describe fundamental principles of machine learning on Azure questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Describe fundamental principles of machine learning on Azure

Choose the correct Azure Machine Learning model type—regression, classification, or clustering—for a scenario, and know how to build it with Azure Machine Learning designer, compute targets, and datasets. The key is matching model type to the prediction task.

Selecting regression, classification, or clustering for a given business scenario

Identifying Azure Machine Learning workspace assets: datasets, compute targets, and environments

Using Automated Machine Learning to train and deploy models with minimal code

Building ML pipelines with Azure Machine Learning designer's drag-and-drop modules

Watch out for

Common Describe fundamental principles of machine learning on Azure exam traps

  • ▸Confusing clustering (unsupervised, no labels) with classification (supervised, labeled categories), which leads to wrong model choice on scenario questions.
  • ▸Assuming Automated Machine Learning removes all data preparation; it still needs clean, correctly formatted tabular data with a defined target column.
  • ▸Mixing up Azure Machine Learning compute targets: compute instances for authoring versus compute clusters for scalable training jobs.

Question index

All Describe fundamental principles of machine learning on Azure questions (199)

Click any question to see the full explanation, or start a practice session above.

1

What is the difference between a binary classification model and a multi-class classification model?

Medium
2

What is 'Bayesian optimisation' in hyperparameter tuning?

Medium
3

A data scientist wants to train a model that predicts whether a customer will respond to a marketing offer (yes or no). The dataset includes features such as age, income, past purchase history, and the labeled outcome (responded or not responded) for previous customers. Which type of machine learning is this?

Easy
4

What does it mean for an ML model to 'generalize'?

Medium
5

A data science team trains a regression model to predict house prices. They evaluate the model using Mean Absolute Error (MAE). After deployment, they notice that the model occasionally produces large errors (e.g., underpredicting a luxury home by $500,000) while most predictions are within $20,000. The business is more concerned about the impact of these large errors than the average small error. Which additional metric should the team use to better capture the penalty for large errors?

Hard
6

A retail company wants to analyze customer purchase histories to identify natural groups of customers with similar buying patterns. They do not have predefined categories. Which type of machine learning should they use?

Medium
7

What is 'model lineage' in Azure Machine Learning?

Medium
8

What is hyperparameter tuning in machine learning?

Medium
9

A retail company wants to predict the exact number of units of a product that will be sold next month. They have historical sales data and information about promotions and holidays. The target variable is the number of units sold, which is a continuous value. Which type of machine learning task should they perform?

Easy
10

A data scientist trains a regression model to predict housing prices. The model uses polynomial features up to degree 5. It achieves an R-squared of 0.95 on the training set but only 0.60 on the test set. Which problem is the model most likely experiencing?

Hard
11

A data scientist trains a deep neural network on a small dataset. The model achieves 100% accuracy on the training data but only 60% accuracy on a validation set. Which technique is most appropriate to address this issue?

Medium
12

What is 'training data' vs 'test data' in machine learning?

Easy
13

A data scientist trains a model to predict house prices. The model achieves 99% accuracy on the training data but only 80% accuracy on new test data. Which technique is most likely to help improve the model's generalization?

Medium
14

What is 'model versioning' and why is it essential in MLOps?

Medium
15

What is 'gradient boosting' and how does it differ from random forests?

Hard
16

A data scientist trains a machine learning model on historical sales data to predict future sales volume. The model achieves 99% accuracy on the training dataset but only 75% accuracy on a separate test dataset. What is the most likely issue with this model?

Medium
17

A robotics team is training a robot to navigate a maze. The robot receives a positive reward (+10) when it reaches the exit and a negative reward (-1) every time it bumps into a wall. The robot learns to maximize its cumulative reward over multiple trials. Which type of machine learning is being used?

Medium
18

What is 'ensemble learning' in machine learning?

Medium
19

A data scientist trains a classification model to predict whether an email is 'phishing' or 'legitimate'. The model achieves 99% accuracy on the training data but only 68% accuracy on the test data. Which action is most likely to help improve the model's generalization performance?

Medium
20

What is 'Azure Machine Learning workspace' and what does it contain?

Easy
21

What does 'deep learning' refer to in machine learning?

Easy
22

A data scientist has a small dataset with only 200 labeled samples. They want to get a reliable estimate of model performance without using a separate validation set that would reduce the training data. Which technique should the data scientist use in Azure Machine Learning to obtain this reliable estimate?

Hard
23

What is 'cross-validation' and when should it be used in machine learning?

Medium
24

What is the purpose of splitting data into training, validation, and test sets in machine learning?

Medium
25

What is 'ensemble learning' in machine learning and why does it improve performance?

Medium
26

A data scientist trains a binary classification model to detect fraudulent transactions. The dataset contains 99% legitimate transactions (negative class) and 1% fraudulent transactions (positive class). The model predicts 'legitimate' for every transaction in the test set and achieves 99% accuracy. Which metric would best reveal that the model is failing to identify any fraudulent transactions?

Medium
27

Drag and drop the steps to train a custom vision model in Azure Custom Vision into the correct order.

Medium
28

A data scientist at a retail company has a dataset of customer transactions with features such as age, income, and purchase history. The goal is to group customers into distinct segments based on similarities in these features, without any predefined labels. Which type of machine learning should the data scientist use?

Medium
29

A data scientist trains a linear regression model to predict house prices. The model's training error is very high, and its test error is nearly as high. Which term best describes this situation?

Medium
30

What is 'differential privacy' and how is it relevant to AI model training?

Hard
31

A data scientist is training a binary classification model to detect fraudulent transactions. The dataset contains only 1% fraudulent transactions. The model achieves 99% accuracy on the test set, but when deployed, it fails to detect most actual fraud cases. Which metric would best reveal this issue?

Hard
32

What is 'time series forecasting' and what Azure ML tools support it?

Medium
33

A data scientist evaluates a regression model that predicts house prices. On the test set, the Mean Absolute Error (MAE) is $8,000 and the Root Mean Squared Error (RMSE) is $25,000. What does the large difference between MAE and RMSE indicate about the model's errors?

Hard
34

A data scientist is building a classification model to predict customer churn. The dataset has only 5% churn cases. The model achieves 95% accuracy on the test set, but upon investigation, the data scientist finds the model predicts 'not churn' for nearly every customer. Which metric should the data scientist primarily use to evaluate the model's performance on this imbalanced dataset?

Medium
35

A data scientist trains a multiclass classification model to identify different species of flowers (Iris setosa, Iris virginica, Iris versicolor). The overall accuracy is 94%, but the accuracy for the Iris virginica class is only 60%. Which additional metric should the data scientist examine to better understand the model's performance on the minority class?

Medium
36

A retail company wants to automatically group customers into segments based on their purchasing history, age, and location without using any predefined labels. The goal is to identify distinct customer profiles for targeted marketing campaigns. Which type of machine learning approach should they use?

Easy
37

A data scientist trains a regression model to predict house prices. The model achieves very low error on the training data but significantly higher error on a held-out test set. Which problem does this scenario best describe?

Medium
38

What is 'hyperparameter tuning' in Azure Machine Learning?

Medium
39

A robotics company is training a drone to fly autonomously through an obstacle course. The drone receives positive rewards for staying on course and avoiding obstacles, and negative rewards for collisions. The system learns by trial and error to maximize its cumulative reward. Which type of machine learning is being used?

Easy
40

A city's traffic department wants to predict the number of cars that will cross a particular bridge each day to plan maintenance schedules. The output of the model should be a numerical value representing the estimated traffic count. Which type of machine learning task is this?

Medium
41

What is 'Azure Machine Learning pipelines' and why are they used?

Medium
42

A data scientist has a dataset containing information about houses: size (sq ft), number of bedrooms, location, and the actual sale price. The goal is to train a model that predicts the price of a new house based on these features. Which type of machine learning task is this?

Medium
43

What is 'model registry' in Azure Machine Learning?

Medium
44

Which Azure service provides a no-code/low-code drag-and-drop interface for building machine learning pipelines?

Medium
45

A data scientist trains a classification model to predict whether an email is spam or not. The model achieves 98% accuracy on the test set, but upon inspection, it classifies all emails as 'not spam' because the dataset has 95% non-spam emails. What is the most likely issue?

Medium
46

Drag and drop the steps to analyze an image with Azure Computer Vision into the correct order.

Medium
47

A data scientist trains a classification model on a dataset of 10,000 labeled emails to distinguish spam from non-spam. The model achieves 99% accuracy on the training data but only 70% accuracy on a held-out test set. Which term best describes this situation?

Medium
48

A real estate company has a dataset containing square footage, number of bedrooms, and location for 10,000 houses, along with their sale prices. They want to train a model that predicts the sale price of a new house based on these features. Which type of machine learning should they use?

Easy
49

What is the role of a label (also called target or ground truth) in supervised machine learning?

Easy
50

What is 'dimensionality reduction' and why is it useful in machine learning?

Easy
51

What is the bias-variance tradeoff in machine learning?

Hard
52

A real estate company trains a model to predict house prices. They evaluate it on a test set of 100 houses. The model predictions have a mean absolute error (MAE) of $5,000 and a root mean squared error (RMSE) of $20,000. What does the large difference between MAE and RMSE indicate about the model's errors?

Hard
53

A data scientist is building a classification model to detect fraudulent transactions. The dataset has 1,000,000 legitimate transactions and only 1,000 fraudulent ones. The model achieves 99.9% accuracy on the test set, but it fails to catch most fraudulent cases. Which metric should the data scientist prioritize to better evaluate the model's performance on this imbalanced dataset?

Hard
54

A company builds a machine learning model to predict whether a customer will purchase a product. They use a training dataset with 50% purchasers and 50% non-purchasers. The model achieves 90% accuracy on the test set. However, when deployed, the model performs poorly because the actual customer base has only 5% purchasers. What is the most likely cause of this poor performance?

Easy
55

What is the purpose of a confusion matrix in evaluating a classification model?

Medium
56

A data scientist trains a regression model to predict house prices. The model has a mean absolute error (MAE) of $5,000 on the test set. Which statement best interprets this metric?

Medium
57

A data scientist trains a decision tree model to predict customer churn. The model achieves 99% accuracy on the training data but only 80% on the test data. Which concept best explains this performance difference?

Medium
58

What is 'active learning' in Azure Machine Learning data labelling?

Medium
59

A data scientist trains a binary classification model to detect spam emails. The dataset contains 95% legitimate emails (negative class) and 5% spam (positive class). The model predicts all emails as legitimate. The accuracy is 95%, but the model is useless. Which metric would best indicate the model's failure?

Hard
60

What is the purpose of Azure Machine Learning's automated ML (AutoML) feature?

Medium
61

A data scientist uses Azure Machine Learning to train a model that predicts the electricity consumption (in kilowatt-hours) of a building based on features like building age, square footage, and number of occupants. The data scientist wants to evaluate how accurately the model's predictions match the actual consumption values. Which evaluation metric is most appropriate for this regression task?

Easy
62

What is the Azure Machine Learning model registry?

Medium
63

What is 'regularization' in machine learning and why is it used?

Medium
64

What is 'model interpretability' and which Azure tool helps with it?

Medium
65

What is the difference between 'precision' and 'recall' as model evaluation metrics?

Hard
66

A data scientist has trained a binary classification model to detect fraudulent credit card transactions. The dataset contains 99.9% legitimate transactions and only 0.1% fraudulent ones. The model predicts all transactions as legitimate, achieving 99.9% accuracy on the test set. However, the business requires the model to actually catch as many fraudulent transactions as possible. Which metric would best reveal the model's failure to identify fraud?

Medium
67

A retail company has historical data about customers, including age, purchase history, and whether they have churned (yes/no). They want to train a model that predicts if a new customer will churn. Which type of machine learning should they use?

Easy
68

What is an ML pipeline in Azure Machine Learning?

Medium
69

A data scientist is training a model to classify customer reviews as positive, negative, or neutral. The dataset contains 10,000 reviews, but only 500 of them are negative. The data scientist wants to ensure the model performs well on the minority class (negative reviews). Which technique should the data scientist consider to address the class imbalance?

Medium
70

What is 'Azure Machine Learning environments' and why are they important for reproducibility?

Medium
71

A data scientist has a dataset containing thousands of labeled images of cats and dogs. The data scientist wants to train a model that can automatically classify new unlabeled images as either 'cat' or 'dog'. Which type of machine learning should the data scientist use?

Easy
72

What is the 'mean absolute error' (MAE) metric used to evaluate in machine learning?

Easy
73

A data scientist trains a binary classification model to predict loan defaults. The dataset contains 98% non-default cases and only 2% default cases. The model predicts 'non-default' for every instance, achieving 98% accuracy on the test set. Which metric would best reveal that the model fails to identify any actual defaults?

Hard
74

Match each Azure AI service to its associated API or SDK.

Medium
75

An e-commerce company has a dataset of customer purchase histories with no predefined categories. The data analyst wants to identify natural groupings of customers based on their purchasing behavior to target marketing campaigns. Which type of machine learning should the analyst use?

Easy
76

A bank uses a machine learning model to predict credit card fraud. The model's output is a probability score. The business wants to minimize the number of false positives (legitimate transactions incorrectly flagged as fraud) because these cause customer dissatisfaction. At the same time, they must also catch most fraudulent transactions. Which metric should the bank optimize to balance these two goals?

Hard
77

A data scientist has a dataset with 100 features and 10,000 samples. They want to reduce the number of features while retaining as much variance as possible, to improve model training speed and reduce overfitting. Which technique should they use?

Medium
78

What is 'ONNX' and why is it relevant to Azure AI?

Medium
79

A data scientist trains a regression model to predict house prices using features like bedrooms, square footage, and location. The model achieves a low error on the training data but performs significantly worse when used to predict prices in a new city with different property characteristics. Which concept best explains this poor performance?

Medium
80

What is 'semi-supervised learning' and when is it useful?

Medium
81

What is 'overfitting' in machine learning and how does Azure ML help prevent it?

Medium
82

What is a training job in Azure Machine Learning?

Medium
83

What is 'model deployment' in Azure Machine Learning?

Easy
84

What is a neural network?

Medium
85

What is 'ROC-AUC' and when is it a better metric than accuracy for classification?

Medium
86

A hospital deploys a machine learning model to screen patients for a rare disease. Only 0.1% of patients actually have the disease. The model correctly identifies most positive cases but also flags many healthy patients as potentially having the disease. The hospital wants to minimize the number of healthy patients who are incorrectly told they might have the disease. Which metric should the model optimize?

Medium
87

A data scientist is preparing a dataset to train a model that predicts customer churn. The dataset includes a column 'CustomerID' which is a unique identifier for each customer. Should the data scientist include the 'CustomerID' column as a feature in the training data?

Easy
88

A data scientist trains a machine learning model to predict house prices based on features like square footage, number of bedrooms, and location. The model achieves a very low error on the training data but performs poorly on a held-out test set. Which term best describes this situation?

Medium
89

Which metric is MOST appropriate for evaluating a regression model's performance?

Medium
90

A data scientist is building a binary classification model to predict fraudulent credit card transactions. The dataset is highly imbalanced: only 1% of transactions are fraudulent. The cost of a false negative is very high because missing a fraudulent transaction can lead to significant financial loss. Which evaluation metric should the data scientist prioritize to minimize false negatives?

Medium
91

A data scientist wants to train a machine learning model to predict the exact market price of a house based on features such as square footage, number of bedrooms, and location. Which type of machine learning task should be used?

Medium
92

A data scientist trains a binary classification model to detect fraudulent credit card transactions. The dataset contains 99.5% legitimate transactions and 0.5% fraudulent transactions. The model predicts every transaction as legitimate and achieves 99.5% accuracy on the test set. Which metric would best reveal that the model is failing to identify any fraudulent transactions?

Medium
93

A data scientist is developing a classification model to detect fraudulent transactions. The dataset is split into training and test sets. The data scientist repeatedly tunes the model's hyperparameters and evaluates performance on the test set until the test accuracy reaches 95%. However, when the model is deployed on new, unseen data, its accuracy drops to 70%. Which concept best explains this performance degradation?

Medium
94

What is an endpoint in Azure Machine Learning?

Easy
95

What does 'model accuracy' measure in machine learning classification?

Medium
96

A medical research team trains a model to detect a rare disease from lab results. The disease occurs in only 1% of patients. The model predicts 'no disease' for every patient and achieves 99% accuracy. Which metric best reveals that the model is failing to identify actual disease cases?

Medium
97

A data scientist trains a model to predict house prices using features like number of bedrooms, square footage, and location. The model achieves a mean absolute error (MAE) of $5,000 on the training data but $25,000 on the test data. Which problem is the model most likely experiencing?

Easy
98

A data scientist trains a model to predict the exact number of cars that will cross a bridge each day for maintenance planning. The model uses historical traffic data as input. Which type of machine learning task is this?

Medium
99

What is 'model evaluation' and what metrics are used for different ML task types?

Medium
100

What is 'regularisation' in machine learning and what problem does it solve?

Medium
101

A data scientist is evaluating a binary classification model that predicts whether a transaction is fraudulent. The test set contains 1,000 transactions: 990 legitimate and 10 fraudulent. The model's predictions are shown in the confusion matrix below. Confusion matrix: Predicted Legitimate Predicted Fraudulent Actual Legitimate 942 48 Actual Fraudulent 2 8 Which metric should the data scientist prioritize if the business goal is to minimize the number of fraudulent transactions that are missed (false negatives)?

Hard
102

A data scientist has a dataset containing customer transaction records with features such as age, income, and purchase history, but no labels. The goal is to identify natural groupings of customers for a targeted marketing campaign. Which type of machine learning should be used?

Medium
103

What is 'Azure Machine Learning Responsible AI dashboard's error analysis'?

Easy
104

What is 'imbalanced classification' handling using 'SMOTE'?

Medium
105

What is a 'compute instance' in Azure Machine Learning?

Medium
106

A data scientist is training a regression model to predict house prices using features like square footage, number of bedrooms, and location. After evaluating the model on a test set, the data scientist wants to select a metric that measures the average magnitude of prediction errors in the same units as the target variable (price). Which evaluation metric should the data scientist use?

Medium
107

What is 'online learning' (incremental learning) in machine learning?

Medium
108

A data scientist is training a model to predict whether a patient has a rare disease (1% prevalence). The model predicts 'no disease' for all patients and achieves 99% accuracy, but fails to identify any actual cases. Which metric would best reveal this failure?

Medium
109

What is 'Azure Machine Learning compute' and what types are available?

Easy
110

A data scientist is training a regression model to predict house prices. The data scientist wants to evaluate the model using a metric that penalizes large prediction errors significantly more than small errors. Which evaluation metric should the data scientist choose?

Medium
111

Which type of machine learning uses labeled training data where the correct output is provided for each input?

Medium
112

A data scientist is building a model to predict the exact temperature in degrees Celsius based on humidity and atmospheric pressure. The model will output a single numeric value for each input. Which type of machine learning task is this?

Easy
113

What is 'transfer learning' and how is it different from training from scratch?

Medium
114

What is 'k-fold cross-validation' specifically and how is k=10 different from k=5?

Medium
115

A data scientist trains a binary classification model to detect fraudulent transactions. The dataset contains only 2% fraudulent transactions. The model achieves 98% overall accuracy, but it fails to detect any fraudulent transactions, classifying all transactions as legitimate. Which metric would most clearly reveal this failure?

Hard
116

A data scientist is training a binary classification model to detect rare equipment failures from sensor data. The dataset contains 99.5% normal operation readings and only 0.5% failure readings. The model currently predicts all readings as 'normal' and achieves 99.5% accuracy on the test set. The business requires the model to identify at least 80% of actual failures. Which data-level technique should the data scientist use to most directly address the class imbalance?

Hard
117

What is feature engineering in machine learning?

Easy
118

What is the F1 score in machine learning evaluation?

Medium
119

What is overfitting in machine learning?

Medium
120

A data scientist wants to group customers into segments based on purchasing behavior without using any labeled examples. Which type of machine learning is this?

Easy
121

A data scientist trains a model on historical data and achieves high accuracy on both the training set and a held-out test set. However, when the model is deployed in production, it performs poorly on new, unseen data. Which issue is most likely the cause?

Medium
122

A hospital has a dataset with historical patient records, each labeled as either 'readmitted within 30 days' or 'not readmitted'. The hospital wants to train a model to predict which current patients are likely to be readmitted. Which type of machine learning task is this?

Medium
123

What is 'Azure Machine Learning's Responsible AI dashboard' and what does it include?

Medium
124

What is 'automated machine learning' (AutoML) in Azure Machine Learning?

Easy
125

What is 'causal inference' and how does it differ from correlation-based machine learning?

Hard
126

What is 'batch inference' vs 'real-time inference' in Azure Machine Learning?

Medium
127

What is 'Azure Machine Learning notebooks' and who typically uses them?

Easy
128

What is data drift in the context of deployed machine learning models?

Medium
129

A media company wants to automatically organize a large collection of news articles into several topic-based categories (e.g., politics, sports, technology) without using any predefined labels. They plan to use Azure Machine Learning. Which type of machine learning task should they use?

Medium
130

A data scientist trains a regression model on a dataset with 100 features and 10,000 samples. The model achieves a low training error but a much higher error on a held-out test set. Which approach is most likely to improve the model's generalization performance?

Medium
131

A data scientist trains a machine learning model on a dataset of housing prices. The model achieves 98% accuracy on the training data but only 72% accuracy on a separate test set. What is the most likely problem with this model?

Medium
132

A data scientist trains a binary classification model to distinguish between images of cats and dogs. On the test set, the model achieves 98% accuracy, but a deeper inspection reveals that the test set contains 95% cats and 5% dogs, and the model predicts 'cat' for every single image. Which metric should the data scientist prioritize to get a more realistic evaluation of the model's performance on this imbalanced dataset?

Easy
133

A data scientist has a dataset containing images of handwritten digits (0-9) where each image is labeled with the correct digit. The goal is to train a model that can predict the digit from a new image. Which type of machine learning approach should be used?

Medium
134

A data scientist trains a regression model to predict house prices. The model performs poorly on both the training data and the test data, showing high error in both sets. Which concept best describes this situation?

Medium
135

What is 'Azure Machine Learning datasets' and why are they important?

Medium
136

A data scientist is training a regression model to predict house prices in Azure Machine Learning. The model uses features like square footage, number of bedrooms, and location (zip code). The data scientist notices that the model has a very low error on the training data but a high error on the test data. Which technique should the data scientist apply during model training to reduce overfitting by penalizing large coefficients?

Medium
137

What is 'stochastic gradient descent' (SGD) and how does it work?

Medium
138

A data scientist trains a regression model to predict daily electricity consumption (in kWh) for a commercial building. The business team needs a metric that heavily penalizes large prediction errors (outliers) more than small errors. Which metric should the data scientist report to best meet this requirement?

Medium
139

A data scientist is training a classification model on a dataset with 100 features and only 500 labeled samples. The model achieves 99% accuracy on the training data but only 68% accuracy on a held-out test set, indicating overfitting. Which technique is most appropriate to directly address this problem?

Medium
140

What is 'label imbalance' in a classification dataset and how does it affect model training?

Medium
141

A data scientist is using Azure Automated Machine Learning to build a binary classification model for a highly imbalanced dataset (95% negative, 5% positive). The data scientist wants AutoML to select the best model based on a metric that is robust to class imbalance. Which primary metric should the data scientist configure in the AutoML settings?

Hard
142

What is model monitoring in Azure Machine Learning and why is it important?

Medium
143

What is 'model monitoring' in Azure Machine Learning after deployment?

Medium
144

What is 'neural architecture search' (NAS) and how does it relate to AutoML?

Hard
145

A bike-sharing company wants to predict the number of rentals per hour. Their model's predictions are usually close but occasionally have large errors due to unexpected events like sudden rain. They want a metric that heavily penalizes these large errors to ensure the model is not overly confident. Which evaluation metric should they primarily use?

Medium
146

What is 'Azure ML's experiment tracking' and why do data scientists use it?

Easy
147

What is machine learning?

Easy
148

A manufacturer trains a model to detect defective parts on an assembly line. Only 2% of parts are defective. The model predicts 'non-defective' for all parts and achieves 98% accuracy. Which metric best reveals the model's inability to identify defective parts?

Medium
149

What is 'curriculum learning' and how does it relate to training stability?

Hard
150

What is 'data augmentation' and how does it help with limited training data?

Medium
151

What is 'clustering' in unsupervised machine learning?

Easy
152

What is 'regression' in machine learning and when is it used?

Easy
153

A manufacturing team wants to predict product defects based on sensor readings from the production line. They have 10,000 historical samples, each labeled as 'defective' or 'non-defective'. Which type of machine learning should they use in Azure Machine Learning?

Medium
154

What is 'data preprocessing' and why is it important for machine learning?

Easy
155

What is recall (sensitivity) in the context of binary classification model evaluation?

Medium
156

What is 'model explainability' using SHAP values in Azure Machine Learning?

Medium
157

What is 'model compression' and what techniques does it include?

Medium
158

What is Azure Machine Learning?

Easy
159

A data scientist is training a logistic regression model to predict customer churn using a small dataset with 500 records and 200 features. The model achieves 97% accuracy on the training set but only 65% on a held-out test set, indicating severe overfitting. The data scientist wants to reduce overfitting by automatically eliminating irrelevant features. Which technique should the data scientist apply?

Hard
160

A retail company wants to automatically group its customers into distinct segments based on their purchasing patterns, without having pre-defined categories. The goal is to discover natural groupings in the customer data to tailor marketing campaigns. Which type of machine learning task should the company use?

Easy
161

What is the Azure Machine Learning workspace?

Easy
162

What is 'confusion matrix' and what does it tell you about a classification model?

Medium
163

A data scientist trains a regression model to predict the selling price of houses. After evaluating on a test set, the data scientist wants a metric that measures the average absolute error between predicted and actual prices, expressed in the same units (dollars) as the target variable. Which evaluation metric should the data scientist use?

Medium
164

What is 'feature engineering' and why does it matter for machine learning models?

Medium
165

A data scientist is building a machine learning model to predict the number of daily bike rentals in a city based on weather data and day of the week. The target variable is a continuous integer. Which type of machine learning task is this?

Medium
166

A data science team trains several machine learning models for a regression task. They observe that Model A has low training error and low test error. Model B has low training error but high test error. Model C has high training error and high test error. Which model would most likely benefit from an ensemble technique that averages the predictions of multiple models?

Hard
167

A retail company has a dataset of customer transaction records with no predefined categories. They want to identify natural groupings of customers based on their purchasing behavior to create targeted marketing campaigns. Which type of machine learning should they use in Azure Machine Learning?

Medium
168

A retail company wants to predict which customers are likely to stop using their service. They have a dataset with many customer attributes including age, income, purchase history, website activity, and support interactions. They suspect some features are redundant. Which technique should they use to reduce the number of features while preserving as much information as possible?

Medium
169

What is Azure Machine Learning's 'responsible AI dashboard'?

Easy
170

What is the role of a validation dataset in machine learning?

Medium
171

A data scientist trains a binary classification model to detect fraudulent transactions. The dataset contains only 1% fraudulent cases. The model predicts 'not fraudulent' for all transactions and achieves 99% accuracy. Which metric would best reveal the model's poor performance on fraud detection?

Medium
172

A data scientist has trained a binary classification model to predict whether an email is spam (positive) or not spam (negative). On a test set, the model correctly identifies 90 out of 100 actual spam emails and 80 out of 100 actual non-spam emails. Which metric shows the proportion of actual spam emails that the model correctly predicted?

Medium
173

What is a training dataset in machine learning?

Easy
174

A data scientist is training a regression model to predict energy consumption. The dataset includes features like temperature, humidity, time of day, and day of week. After training, the model performs well on the training set but poorly on new data. Which approach would most likely help reduce this problem?

Medium
175

What is 'feature importance' in Azure Machine Learning and how is it used?

Medium
176

What is a feature in the context of machine learning?

Medium
177

What does Azure Machine Learning's 'compute cluster' provide?

Easy
178

A retail company wants to segment its customers into different groups based on purchasing behavior, without using predefined categories. Which type of machine learning task should they use?

Medium
179

A data scientist is building a machine learning model to predict whether a credit card transaction is fraudulent or legitimate. The dataset contains 100,000 historical transactions, each labeled as 'fraudulent' or 'legitimate'. Which type of machine learning task should the data scientist use in Azure Machine Learning?

Medium
180

What type of machine learning model is used for time series forecasting?

Easy
181

What is the purpose of a 'validation dataset' in machine learning?

Easy
182

What is 'Azure Machine Learning's job submission' and what types of training jobs are supported?

Medium
183

An online retailer wants to build a recommendation system that learns from user interactions. The system suggests a product, and if the user clicks it, it receives a positive reward; if ignored, a negative reward. Over time, the system learns to make better suggestions. Which type of machine learning best describes this approach?

Hard
184

What is the purpose of Azure Machine Learning's dataset versioning?

Medium
185

What is the difference between 'training' and 'inference' in machine learning?

Easy
186

What is a confusion matrix's 'false positive' in medical screening?

Medium
187

What is 'federated learning' and when is it used for privacy-preserving AI?

Hard
188

What is cross-validation in machine learning?

Medium
189

A data scientist trains a regression model to predict energy consumption for a smart building. The model achieves very low error on the training data but performs significantly worse on a held-out validation set. Which technique would most directly address this problem?

Hard
190

A data scientist trains a machine learning model to predict housing prices. On the training data, the model achieves an R-squared value of 0.99, but on a separate validation dataset it achieves an R-squared of only 0.65. What is the most likely issue with this model?

Easy
191

A data scientist trains a binary classification model to detect a rare disease. The dataset contains 99% negative cases and only 1% positive cases. The model predicts all cases as negative, achieving an accuracy of 99% on the test set. However, the business requires the model to identify as many positive cases as possible. Which metric should the data scientist examine to best reveal that the model is failing to identify any positive cases?

Medium
192

What is 'Azure Machine Learning designer' and who is it designed for?

Easy
193

What is reinforcement learning?

Medium
194

A data scientist is building a machine learning model to predict whether a loan applicant will default. The dataset includes features such as credit score, income, and loan amount, and the target variable is 'default' (yes/no). The data scientist wants to evaluate the model's performance using a metric that balances both false positives and false negatives. Which metric should the data scientist use?

Medium
195

What is AutoML in Azure Machine Learning and what does it automate?

Easy
196

What is the purpose of a test dataset in machine learning model development?

Easy
197

A data scientist trains a classification model to distinguish between images of cats and dogs. The model achieves 99% accuracy on the training set but only 75% accuracy on a validation set. Which concept best describes this situation?

Easy
198

What is precision in the context of binary classification model evaluation?

Medium
199

A data scientist is training a model to predict whether a customer will purchase a product (Yes/No). The dataset contains 90% 'No' and 10% 'Yes'. After training, the model achieves 90% accuracy. Which evaluation metric would be more informative to assess the model's performance on the minority class?

Medium

Frequently asked questions

What does the Describe fundamental principles of machine learning on Azure domain cover on the AI-900 exam?
Choose the correct Azure Machine Learning model type—regression, classification, or clustering—for a scenario, and know how to build it with Azure Machine Learning designer, compute targets, and datasets. The key is matching model type to the prediction task.
How many questions are in this domain?
This page lists all 199 Describe fundamental principles of machine learning on Azure questions in the AI-900 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Describe fundamental principles of machine learning on Azure questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
ai-900 AI-900 describe fundamental principles of machine learning on azure Practice Questions