You must be able to prepare and govern data for Salesforce Einstein features. The single most important thing is matching the correct data action to the feature: investigate important fields, secure consent before training, supply required components, and handle missing values without introducing bias.
Start practicing
Data for AI — choose a session length
Free · No account required
Domain overview
This domain covers how Salesforce AI features consume, prepare, and govern data. It is tested through scenario questions about Einstein Discovery, Prediction Builder, Article Recommendations, and Recommendation Builder, asking you to pick the correct data preparation or privacy action rather than define a term. Expect questions on missing values, field importance, consent, and required data components.
Exam objectives
Selecting the right action after Einstein Discovery reports high field importance, such as investigating Case_Origin__c.
Identifying data privacy steps, like consent and lawful basis, before training a recommendation model on purchase history.
Naming required data components for Einstein Article Recommendations, including knowledge articles and case data.
Choosing best practices for missing values in Einstein Prediction Builder datasets to reduce model bias.
Treating high field importance as a defect to remove instead of a signal to investigate the underlying business process.
Assuming any historical purchase data can train a model without checking consent, purpose limitation, or retention rules.
Filling missing numeric values with zero or the mean without considering whether that skews the prediction and introduces bias.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A company wants to use Einstein Prediction Builder to predict customer churn. Which data preparation step is essential before building the model?
2A data scientist needs to prepare data for Einstein Discovery. The dataset includes a field 'Customer_Status__c' with values 'Active', 'Inactive', and 'Churned'. How should this field be treated?
3A Salesforce admin is training an Einstein Bot to answer customer questions. Which data source should the bot use to provide accurate responses?
4A company uses Einstein Discovery to identify factors that increase case resolution time. After training, the model shows that 'Case_Origin__c' has high importance. What action should the company take?
5A company has set up Einstein Next Best Action with a recommendation strategy. They want to ensure that recommendations are personalized based on the customer's recent behavior. What data should be used?
6Which TWO actions are required to prepare data for an Einstein Discovery model?
7A company uses Salesforce Data Platform to store customer data. They want to use this data to train an AI model for lead scoring, but they are concerned about data quality. Which step should they take first to ensure the data is suitable for AI?
8A data scientist is building a predictive model for customer churn using Salesforce data. The dataset has 20 features, and the target variable is highly imbalanced (5% churn, 95% non-churn). Which technique should be applied to handle the class imbalance before training?
9An administrator is configuring a Salesforce AI model that uses historical sales data. The data includes fields like 'Amount', 'Close_Date', and 'Lead_Source'. What is the primary purpose of data preprocessing in this context?
10A company is deploying an AI model that recommends next best actions for sales reps. They notice that the model's recommendations are biased towards high-revenue opportunities. Which data-related action can help reduce this bias?
11A Salesforce admin wants to use Einstein Prediction Builder to predict case resolution time. What type of data is most critical for training this model?
12Which TWO techniques are commonly used to handle missing values in a dataset for AI training?
13Which THREE factors should be considered when selecting features for a predictive model in Salesforce?
14Which TWO are common data quality issues that can negatively impact AI model performance?
15A company wants to use Einstein Article Recommendations to surface relevant knowledge articles to its support agents. What two data components are required to set up this feature?
16A company is using Einstein Discovery to predict customer churn. The model was created six months ago and has been making predictions. Recently, the model's accuracy has dropped significantly. The data scientist confirms that the data schema has not changed. What is the most likely reason for the drop in accuracy?
17A company is implementing Einstein Prediction Builder to predict whether a support case will escalate. Which TWO data preparation steps should the admin take to improve model accuracy?
18You are a Salesforce AI Specialist at a mid-sized manufacturing company. The company uses Einstein Lead Scoring to prioritize leads. The model was trained on historical lead data and has been in production for three months. Recently, the sales team reports that high-scoring leads are not converting as expected. You investigate and find that the model's data source includes leads from the past 18 months. However, six months ago, the company changed its lead qualification process: they started requiring a demo before scoring leads as 'qualified.' As a result, the definition of a converted lead changed. What is the best course of action to improve model performance?
19You are an admin at a financial services firm. The firm wants to use Einstein Next Best Action to offer personalized product recommendations to customers on its service portal. The data includes customer profiles, transaction history, and support case history. The Einstein Next Best Action strategy is configured with a recommendation that shows a 'Savings Account' offer to customers who have a checking account. However, the recommendation is not appearing for any customers. You check the Data Flow and see that the 'Account' object data is flowing correctly. The recommendation's filter condition is: AND( Has_Checking_Account__c = true, Age__c > 18 ). You verify that many customers meet these conditions. What is the most likely reason the recommendation is not appearing?
20You are a data scientist at a retail company. The company uses Einstein Discovery to analyze customer purchase patterns. The model is built on a dataset of 50,000 transactions. The model's R-squared is 0.85, but the predictions for new customers are consistently off by a large margin. The data includes features like 'Customer Age', 'Income', 'Previous Purchases', and 'Product Category'. The model was trained on data from the past two years. However, six months ago, the company launched a new loyalty program that significantly changed purchasing behavior. You suspect the model is not generalizing to new customers. What should you do to validate your hypothesis?
21A company is preparing customer data for a predictive model. They notice that many records have missing values for the 'annual income' field. Which approach is best to handle this issue while minimizing bias?
22For a real-time AI application that requires low-latency access to customer interaction data, which storage solution is most appropriate?
23A company wants to use customer purchase history to train a recommendation model. Which action is essential to comply with data privacy regulations?
24Which data transformation is most appropriate for converting categorical variables into numerical format for a machine learning model?
25Which method is most suitable for ingesting streaming data from IoT sensors into a data lake?
26A global company needs to ensure that customer data used for AI models complies with multiple regional regulations (GDPR, CCPA, LGPD). Which data governance practice is most effective?
27Which TWO data preparation steps are critical for ensuring high-quality training data?
28Which TWO considerations are important when labeling data for a supervised learning model?
29A data engineer is troubleshooting a predictive model that stopped updating. The data flow from Data Cloud shows 'Data Transform Failed' with error: 'Field Amount cannot be null'. What is the most likely cause?
30A company is preparing data for Einstein Article Recommendation. Which data source is most appropriate for training the model?
31An admin created a data stream to bring external customer data into Data Cloud for Einstein. The data stream fails with error 'Schema mismatch: expected 10 fields, got 8'. What is the likely cause?
32A data architect is designing a data model for Einstein Discovery. The data includes categorical variables with high cardinality (e.g., postal codes). What is the best practice to handle such features?
33A system administrator receives an error when running a Data Cloud data transform: 'Row-level security settings are preventing access to the source data.' The admin has appropriate permissions. What is the most likely cause?
34A marketer wants to use Einstein Segment Creation to build a segment for a campaign. Which data source can be used?
35A data analyst is evaluating data quality for an Einstein model. Which TWO dimensions are most critical for model accuracy?
36A company is ingesting data from multiple sources into Data Cloud for Einstein. Which THREE data preparation steps should be performed?
37A company is preparing data for Einstein Prediction Builder to forecast lead conversion. They have historical data with fields like Lead Source, Industry, Number of Employees, and Converted (boolean). Which data preparation step is most critical?
38A data scientist notices that an Einstein model for predicting customer churn has unusually high accuracy on training data but performs poorly on validation data. Which data issue is the most likely cause?
39A company wants to build a sentiment analysis model using customer feedback. What is the best practice for labeling the training data?
40A large enterprise needs to integrate data from Salesforce CRM, an external ERP, and marketing automation to train an AI model for cross-sell recommendations. Which data storage strategy is most aligned with Salesforce's AI capabilities?
41A company is using customer support tickets to train a model for auto-classifying issues. The dataset includes fields like 'Case Title', 'Description', 'Product', and 'Customer Name'. Which privacy concern is most critical to address before training?
42A fraud detection model is being trained on transaction data where only 1% of transactions are fraudulent. The current model predicts 'non-fraud' for all transactions, achieving 99% accuracy. Which technique should be applied to improve model performance?
43After applying a log transformation to a numeric feature, an Einstein model’s performance dropped significantly. What is the most likely cause?
44A bank uses Einstein Discovery to generate insights about loan approval decisions. After deployment, they notice the model denies loans to a higher percentage of applicants from a certain postal code. Which action should be taken to ensure responsible AI?
45A company plans to use Einstein Discovery to analyze sales data. Which data preparation step is essential for time-series forecasting?
46A data scientist is preparing numeric features for a regression model. Which TWO transformations are commonly applied to improve model performance?
47Refer to the exhibit. A data analyst has defined this field mapping for Einstein Prediction Builder. Which data issue would most likely arise from this mapping?
48A data scientist notices that the model accuracy drops significantly after retraining with new data. Upon inspection, they find that many records have missing values for a key feature. Which data quality improvement should be prioritized first?
49A company is building a text classification model for customer support tickets. They have a dataset of 10,000 tickets. The team decides to use active learning for labeling. Which approach best aligns with active learning principles?
50A team is building a pipeline to train a model daily. The source data arrives in CSV files but needs to be converted to Parquet for efficiency. Which pipeline step should perform this conversion?
51During data transformation, a data scientist applies one-hot encoding to a categorical feature with 50 unique values. The resulting dataset has 50 new columns. What is a potential drawback of this transformation?
52An organization uses Salesforce Data Cloud to unify customer data from multiple sources. They want to ensure that data lineage is tracked for AI models. Which practice supports data lineage?
53Which TWO of the following are common dimensions of data quality that must be addressed for AI training?
54Which TWO considerations are critical when planning data labeling for a computer vision project in a regulated industry?
55Which THREE types of data sources are commonly integrated into Salesforce Data Cloud for AI use cases?
56Refer to the exhibit. A data pipeline fails during the DataTransformation stage. What is the most likely root cause?
57Refer to the exhibit. A data transformation configuration is shown. Which of the following describes the outcome of applying this transformation?
58A company uses Einstein Prediction Builder to predict customer churn. The data includes account creation date, number of support cases, and average payment delay. After training, the model shows low confidence scores. What is the most likely cause?
59A Salesforce admin wants to use Einstein Recommendations to suggest products. What is a key requirement for the data used to train the recommendation model?
60An admin is setting up Einstein Article Recommendations. Which type of data is essential for the model to learn which articles are relevant?
61An admin is troubleshooting Einstein Sentiment. The model returns high confidence but wrong sentiment (e.g., positive reviews labeled negative). What is the most likely issue?
62When using Einstein Lead Scoring, which data source is most critical for generating accurate lead scores?
63A company has international customers and wants Einstein Prediction Builder to forecast deal closure probability. The data includes fields like 'region', 'product line', and 'deal amount'. What is a best practice to ensure the model works for all regions?
64A data analyst is troubleshooting Einstein Article Recommendations that are not showing up on the site. Which TWO checks should be performed first? (Choose 2)
65A data scientist needs to feed customer interaction data into Einstein Discovery for predictive analysis. Which data format is required?
66A company uses Salesforce Data Cloud to unify customer data from multiple sources for AI model training. After adding a new data source, model performance degrades significantly. What is the most likely cause?
67Which data type is most commonly used for image recognition AI models?
68A team has limited labeled data for a Salesforce predictive model but wants to leverage a pre-trained model from a related task. Which machine learning approach should they use?
69After deploying an AI model in Salesforce, the data scientist notices high accuracy on the training set but poor accuracy on new incoming data. What is this phenomenon called?
70To ensure AI model fairness and avoid biased outcomes, which practice is most critical when preparing training data?
71A company wants to integrate external customer behavior data into Salesforce to enhance AI predictions. Which Salesforce Data Cloud feature is specifically designed to ingest and map external data?
72A data scientist discovers that an AI model used for loan approval predicts high default risk disproportionately for a specific demographic group. What is the first step to address this issue?
73Which TWO are best practices for data labeling in AI projects? (Choose two.)
74Which THREE are key considerations for data privacy when using AI models that process customer data? (Choose three.)
75Which TWO are common data quality issues that negatively impact AI model performance? (Choose two.)
76What is the primary purpose of this policy?
77A Salesforce admin is preparing a dataset for Einstein Prediction Builder. The dataset contains a field "Income" with many missing values. The admin wants to minimize bias in the model. What is the best practice?
78When training an Einstein Discovery model, which data type is not supported as a predictor field?
79A data scientist notices that a Salesforce Einstein model's performance degrades over time. The model was trained on data from the last year. What is the most likely cause?
80To integrate external data into Salesforce for AI, which tool is recommended by Salesforce for building data pipelines?
81In Salesforce CRM Analytics (formerly Einstein Analytics), what is the primary purpose of a dataset?
82A company wants to use Einstein Next Best Action but needs to ensure data privacy. What is the required step for anonymizing customer data in Data Pipelines?
83While building a prediction model in Einstein Studio, the system warns about "high cardinality" for a categorical field. What should the admin do?
84Which Salesforce feature automatically flags data quality issues before training an AI model?
85A data integration specialist is using Data Pipelines to combine Salesforce data with an external CSV file. The CSV has a header row but some rows have extra commas, causing parsing errors. What should the specialist do?
86A Salesforce admin is reviewing data sources for Einstein Recommendation Builder. Which two data types are required for training? (Choose two.)
87Which three practices help maintain data quality for AI models in Salesforce? (Choose three.)
88When preparing data for Einstein Next Best Action, which two aspects must be considered for compliance with data privacy regulations? (Choose two.)
89Refer to the exhibit. In the JSON configuration above, which data preparation step could introduce bias?
90Refer to the exhibit. What is the most likely cause of the pipeline failure?
91A global company uses Salesforce Einstein Discovery to predict customer churn. They have a dataset with fields: Customer_Since__c (date), Last_Interaction_Date__c (date), Support_Cases__c (number), Product_Usage__c (percentage), Region__c (picklist), and Churned__c (boolean target). The model was trained and deployed, but predictions show bias against customers in the "EMEA" region. The data scientist notices that in the training data, 80% of EMEA customers are labeled as churned, while only 20% of other regions. Additionally, the Product_Usage__c field has many missing values for EMEA customers. The company wants to retrain the model to reduce bias. What is the best course of action?
92A marketing agency needs to ingest real-time social media mentions for a sentiment analysis AI model. Which Data Cloud object type should they use to set up the ingestion?
93A retailer's AI model for recommendation is producing poor results. Analysis shows that the customer entity has many duplicate records with slight variations. Which Data Cloud feature should be used to address this?
94A large enterprise uses Data Cloud to power an Einstein model for lead scoring. The model's feature pipeline includes dozens of fields from multiple data streams. Performance has degraded, and the team suspects slow feature retrieval. What is the most efficient way to speed up feature computation in Data Cloud?
95A financial institution must ensure that customer data used for AI models does not expose personally identifiable information (PII) to unauthorized users. Which Data Cloud feature should be applied to the data model?
96A data architect notices that a Data Stream from an external ERP system is failing intermittently with schema mismatch errors. The ERP team says the schema changes occasionally. What is the most effective long-term solution?
97A news outlet wants to build an AI model that predicts article popularity using real-time social media mentions. Which data source type should they use to ingest tweets?
98A manufacturer wants to improve demand forecasting by enriching its CRM orders with external demographic data. The external data is available via a SOAP API. How should the data architect implement this?
99Which TWO of the following are valid methods to improve data quality in Data Cloud before training an AI model?
100Which THREE of the following are best practices for feature engineering in Einstein Studio?
101A financial services firm uses Data Cloud to enrich sales data with external credit scores via an API. They set up a Data Action to call the credit bureau API for each new lead. Over time, API costs are rising, and the action is slowing down lead processing. They only need credit scores for leads with a high probability of conversion. What is the best approach to reduce costs and improve performance?
102A non-profit organization uses Data Cloud to manage donor data from multiple sources (email campaigns, event attendance, donations). They want to use an AI model to predict future donations. The data scientist says the model needs a unified view of each donor with consistent fields. What is the first step the data architect should take in Data Cloud to enable this?
103A company is preparing customer data to train a custom AI model for sentiment analysis. Which two data preparation best practices should they follow? (Choose two.)
104A retail company has implemented a Salesforce AI lead scoring model to prioritize high-value customers. After three months, the model's AUC-ROC score is only 0.55, indicating poor performance. The data scientist reviews the training data and finds that 20% of the records are exact duplicates due to multiple data imports from different sources. The duplicates have inconsistent target labels (some labeled 'converted', others 'not converted'). What should the data scientist do to improve model performance?
105A telecom company uses Einstein Discovery to predict customer churn. The training dataset contains 100,000 records, but only 5% represent churned customers. The model achieves 95% accuracy on a holdout test set, but the recall for churn is only 20%. The business wants to proactively retain at-risk customers, so they need to identify as many churners as possible. What action should the data scientist take to improve churn recall?
106A healthcare organization uses Salesforce to develop an AI model for patient readmission prediction. They must comply with HIPAA regulations. The dataset includes patient names, addresses, medical record numbers, and detailed clinical notes. The data scientist plans to train a supervised model using historical readmission outcomes. What is the most important data governance step before model training?
107A company is building a chatbot using Einstein Bot's AI capabilities. They want to train intent recognition using historical chat transcripts. The transcripts contain many typos (e.g., 'hellp' instead of 'help') and slang (e.g., 'gonna' instead of 'going to'). The initial model performs poorly, misclassifying many intents. What data cleaning step is most important?
108A multinational corporation uses Salesforce AI to analyze customer feedback across multiple languages. They have 10,000 English reviews, 2,000 Spanish reviews, and 500 French reviews. The sentiment model performs well on English (F1=0.85) but poorly on French (F1=0.40). The data scientist wants to improve French sentiment performance without collecting new data. What should they do?
109Refer to the exhibit. A data analyst receives an error when trying to use this model configuration for Einstein AI predictions. Which issue is most likely causing the error?
110A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts customer churn. They notice that the 'Customer_Email__c' field contains multiple blank values across thousands of records. Which data preparation step should they take FIRST to address this issue?
111A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts customer churn. The dataset contains a 'LastPurchaseDate' field, and the admin wants to transform it into a numeric feature representing the number of days since the last purchase relative to today. Which approach should the admin use to create this feature?
112A data engineer is building a dataset for an Einstein Discovery model that predicts opportunity win probability. The dataset includes a 'Close_Date__c' field. They want to create a feature that captures the number of days between opportunity creation and close. Which transformation should they apply?
113A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts customer churn. The dataset contains a field named 'CustomerSince' with values like '2021-03-15'. The model treats this field as a categorical variable, which lowers accuracy. What should the admin do to improve the model?
114A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts customer churn. The dataset includes a field named 'CustomerID' that uniquely identifies each customer, plus several numeric and categorical fields. The admin wants to ensure the model learns meaningful patterns without memorizing individual customers. What should the admin do with the 'CustomerID' field?
115A data scientist is preparing data for an Einstein Discovery model to predict customer lifetime value. They have a dataset with 50,000 records and 200 features. They notice that one feature, 'Total_Purchases__c', has a highly skewed distribution with a few extreme outliers. Which data preparation technique should they consider to reduce the impact of outliers?
116A Salesforce administrator is setting up a Data Cloud data stream to ingest customer order data from an external system. The external system sends records with a field 'OrderDate' in the format 'MM/DD/YYYY'. The administrator needs to ensure the data is correctly mapped and usable for AI models. What should the administrator do first?
117A Salesforce administrator is preparing data for an AI model that predicts which leads are most likely to convert. They need to ensure the data is high quality. Which two data quality dimensions are most critical to address before training the model? (Choose two.)
118A data analyst is preparing a dataset for an Einstein Prediction Builder model that predicts whether a customer will make a repeat purchase. The dataset contains a field 'LastPurchaseDate' with some missing values. The analyst wants to handle these missing values appropriately. What is the recommended approach?
119A data engineer is preparing a dataset for an Einstein Discovery model to predict customer churn. The dataset has a field 'TotalSpend' that is highly skewed, with a few customers having extremely high values. The engineer wants to reduce the impact of these outliers on the model. Which data transformation should the engineer apply?
120A Salesforce administrator is configuring a Data Cloud data stream to ingest customer profiles from an external CRM. The source sends a 'LastModifiedDate' field in the format 'MM/DD/YYYY', but Data Cloud expects ISO 8601 ('YYYY-MM-DD'). Which Data Cloud feature should be used to transform the date format during ingestion?
121A data scientist is building an Einstein Discovery model to predict which customers are likely to default on a loan. The dataset has a field 'Income' with many missing values. The scientist decides to impute missing values using the median income. After training, the model performs well on the training data but poorly on new data. What is the most likely cause of this issue?
122A data engineer is preparing a dataset for an Einstein Discovery model to predict loan default risk. The dataset includes a field 'Income' with a skewed distribution and several outliers. The engineer wants to improve model performance by transforming this field. Which transformation is most appropriate?
123A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts customer churn. The dataset contains a field named 'LastPurchaseDate' with values that include future dates (e.g., 2025-12-31) for some customers. What should the admin do to address this data quality issue?
124A Salesforce administrator is preparing data for an AI model that predicts customer lifetime value. The dataset contains a 'Country' field with values like 'USA', 'Canada', and 'Mexico'. The administrator wants to use this field in the model but knows that machine learning algorithms require numeric input. What should the administrator do to make the 'Country' field suitable for the model?
125A data engineer is preparing a dataset for an Einstein Discovery model that predicts customer churn. The dataset contains missing values in several fields. Which two methods are appropriate for handling missing values to maintain model accuracy? (Choose two.)
126A Salesforce admin is preparing a dataset for an AI model that predicts customer satisfaction. The dataset includes a field 'Region' with values like 'North', 'South', 'East', and 'West'. The admin wants to ensure the model can use this field effectively. What should the admin do?
127A Salesforce administrator is setting up a Data Cloud data stream to unify customer profiles from multiple sources for use in AI models. The sources include a CRM, a marketing automation platform, and a legacy system. The administrator needs to ensure that records from these sources are matched correctly to create a unified individual. Which Data Cloud feature should be used?
128A data scientist is preparing a dataset for an Einstein Discovery model that predicts customer satisfaction scores. The dataset includes a field 'Region' with values like 'North', 'South', 'East', 'West', and 'Unknown'. The data scientist notices that 'Unknown' appears in 30% of records. What is the most appropriate action regarding this field?
129A Salesforce admin is setting up a dataset for an Einstein Prediction Builder model to predict whether a lead will convert. The dataset includes a field 'Industry' that has many missing values. What is the recommended approach to handle missing values in this categorical field?
130A data engineer at a retail company is preparing a dataset in Salesforce Data Cloud to train a propensity model. The dataset contains a field named 'Preferred_Contact_Time' with values such as 'Morning', 'Afternoon', 'Evening', and 'Weekend'. The model requires numerical input, but the field is stored as text. The engineer wants to transform this field into a format suitable for the model. Which approach should the engineer take?
131A Salesforce admin is using Einstein Discovery to build a model predicting opportunity win probability. The dataset includes a 'CloseDate' field. The admin notices that the model's performance is poor and suspects data leakage. Which action should be taken to prevent leakage from the 'CloseDate' field?
132A data engineer is preparing a dataset for an Einstein Discovery model to predict customer lifetime value. The dataset includes a field 'TotalSpend' that is highly skewed with a few extreme outliers. Which transformation should the engineer apply to improve model performance?
133A data scientist is preparing a dataset for an Einstein Discovery model to predict customer churn. The dataset includes a 'CustomerSince' date field, a 'TotalSpend' numeric field, and a 'SupportTickets' numeric field. The data scientist wants to create new features that capture customer tenure and engagement. Which two feature engineering techniques should be applied? (Choose two.)
134A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts customer churn. The dataset contains a 'LastPurchaseDate' field with dates in multiple formats (e.g., '2023-01-15', '01/15/2023', '15-Jan-2023'). The admin wants to ensure the field is usable for modeling. What should the admin do first?
135A data scientist is building a propensity-to-buy model in Salesforce using Einstein Discovery. The dataset includes a field 'ProductCategory' with 50 distinct values. After training, the model shows high accuracy but performs poorly on new data. What is the most likely cause and the best action to take?
136A Salesforce administrator is preparing a Data Cloud data stream to feed an Einstein Studio predictive model. The source object contains a 'ContactEmail' field where several records have values like 'not provided' or 'n/a'. The administrator wants the model to ignore these records during training. Which data quality dimension is primarily being addressed?
137A data engineer is building a Data Cloud data model for an Einstein Discovery model that predicts whether a customer will renew a subscription. The source system sends the renewal date in the format MM/DD/YYYY as a text string. The engineer notices that Einstein Discovery is treating this field as a categorical variable instead of a date. What should the engineer do to correct this?
138A data scientist is building an Einstein Discovery model to predict opportunity win rates. The dataset includes a 'Probability' field that is manually entered by sales reps and often left blank or set to default values. The data scientist wants to improve model accuracy. What is the most appropriate action regarding the 'Probability' field?
139A Salesforce administrator is integrating external data into Data Cloud to enhance an AI model. The external data contains customer addresses with inconsistent formatting (e.g., 'St.' vs 'Street', 'Ave' vs 'Avenue'). What is the most effective approach to prepare this data for the model?
140A data engineer at a retail company is building a propensity-to-buy model in Einstein Studio using Data Cloud. The training dataset includes a 'LastPurchaseDate' field. The engineer notices that many records have purchase dates from 2018, while the model is intended to predict purchases for the upcoming holiday season in 2024. Which data quality dimension is most directly compromised?
141A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts case escalation. The dataset includes a field for customer satisfaction score (CSAT), but 40% of the records have no CSAT value because surveys were not returned. Which action should the admin take to address the missing values before training the model?
142A Salesforce administrator is preparing data for an AI model that predicts customer lifetime value. They notice that the 'AnnualRevenue' field on Account records has many outliers, such as negative values and extremely high numbers. What should the administrator do to address this data quality issue?
143A data engineer is integrating external data into Data Cloud for an AI model. The external data contains a 'Gender' field with values 'M', 'F', 'Male', 'Female', and 'Unknown'. Which Data Cloud feature should be used to standardize these values before mapping to the Data Model Object?
144A financial services company uses Data Cloud to create a unified customer profile for an AI model that predicts loan default risk. The data engineer notices that the 'Income' field from the CRM system and the 'AnnualIncome' field from an external data source frequently disagree for the same customer. Which action should the engineer take to best address this data quality issue for the AI model?
145A data engineer is preparing a dataset for an Einstein Discovery model to predict customer churn. The dataset includes a field 'CustomerSince' that records the date a customer first joined. The engineer wants to create a new feature that represents the tenure of the customer in months. Which approach is most appropriate?
146A Salesforce data architect is preparing data for an Einstein Discovery model that predicts opportunity win probability. The source data includes opportunity amounts, close dates, stages, and account industry. The architect wants to ensure the model has high-quality training data. Which two actions should the architect take? (Choose two.)
147A data scientist is building an Einstein Discovery model to predict customer lifetime value (CLV). The dataset includes a field 'TotalPurchaseAmount' that is highly skewed, with a few customers having extremely high values. The scientist wants to reduce the impact of outliers on the model. Which data transformation technique should they apply?
148A data engineer is using Data Cloud to prepare a training dataset for an Einstein Discovery model that predicts customer lifetime value (CLV). The dataset contains transactional records at the line-item level, with multiple rows per customer. The model requires one row per customer. Which Data Cloud feature should the engineer use to aggregate the line items into a single customer-level record?
149A telecommunications company is preparing data for an Einstein Studio churn prediction model. The data engineer notices that the 'CustomerTenure' field has a skewed distribution: most customers have tenure under 2 years, but a few have tenure over 20 years. The engineer wants to improve model performance by addressing this skew. Which data preparation technique should the engineer apply?
150A data engineer is preparing a dataset for an Einstein Discovery model that predicts customer churn. The dataset includes fields such as 'LastPurchaseDate', 'TotalSpend', 'CustomerServiceCalls', and 'Age'. Which two data preparation steps are essential to ensure the model performs well? (Choose two.)
151A Salesforce admin is preparing a dataset for an Einstein Prediction Builder model that predicts whether a lead will convert. The dataset contains several fields, including 'Industry', 'AnnualRevenue', 'NumberOfEmployees', and 'LeadSource'. The admin notices that 'Industry' has many missing values, 'AnnualRevenue' has some outliers, and 'LeadSource' has inconsistent entries (e.g., 'Web', 'Website', 'web'). Which TWO data preparation steps should the admin take to improve data quality for the model? (Choose two.)
152A Salesforce data architect is preparing a Data Cloud dataset to train an Einstein Studio model for predicting customer lifetime value. The dataset includes fields such as 'CustomerID', 'LastPurchaseDate', 'TotalSpend', and 'PreferredChannel'. The architect must ensure the data is ready for AI. Which two actions should the architect take to address potential data quality issues that could degrade model performance? (Choose two.)
153A data scientist is building a predictive model in Salesforce Data Cloud to forecast next quarter's sales. The dataset includes historical sales, marketing spend, and economic indicators. The data scientist notices that the 'Marketing_Spend' feature has a high correlation with the target 'Sales' but also with another feature 'Economic_Indicator'. The model's performance is unstable, with high variance in predictions. Which issue is most likely causing this instability?
154A Salesforce admin is preparing a dataset for an Einstein Discovery model that predicts whether a lead will convert. The dataset contains a 'LeadSource' field with values such as 'Web', 'Phone', 'Email', and 'Webinar'. The admin notices that the field has a high cardinality, with over 100 unique values due to free-text entries. What should the admin do to prepare this field for the model?
155A Salesforce admin is using Data Cloud to create a calculated insight that computes the average order value per customer over the last 90 days. The admin needs to ensure the insight updates daily. Which Data Cloud feature should they use to schedule the calculation?
Deep-dive questions
The most-searched questions in this domain — detailed explanations, worked examples, full answer breakdowns.
You must be able to prepare and govern data for Salesforce Einstein features. The single most important thing is matching the correct data action to the feature: investigate important fields, secure consent before training, supply required components, and handle missing values without introducing bias.
The Courseiva AI Associate question bank contains 155 questions in the Data for AI domain, covering the 36% of the exam attributed to this domain in the official Salesforce blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Data for AI domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included