Explain Model Action Set
Compute the Shapley Values (Kernel SHAP Method) Using Prescored Data
This section contains PROC CAS code.
Note: Input data must be accessible in your CAS session, either as a CAS table or as a transient-scope table. A CAS table has a two-level name: the first level is your CAS engine libref, and the second level is the table name. You refer to this table in the CAS procedure by specifying only the second level. For more information about two-level names, see Chapter 2, Shared Concepts (SAS Viya: Machine Learning Procedures). A transient-scope table is called directly from the action and exists in memory for the duration of the action. For more information about accessing data, see SAS Viya: System Programming Guide. For more information about PROC CAS and programming in CASL, see SAS Cloud Analytic Services: CASL Programmer’s Guide and SAS Cloud Analytic Services: CASL Reference.
This example runs the linearExplainer action to explain a prediction that is made by a forest model by using the Kernel SHAP method with prescored data.
The following DATA step creates the reference data table mycas.dmagecr in your CAS session; adds the ID variable id; and keeps only the variables age, amount, coapp, duration, foreign, good_bad, id, and job. These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.
data mycas.dmagecr;
set sampsio.dmagecr;
id = _N_;
keep age amount coapp duration foreign job good_bad id;
run;
The following code runs the forestTrain action in the decisionTree action set to build a forest model in order to predict whether the credit rating of each individual in the dmagecr data table is good or bad. The action also outputs an analytic store named forest_astore.
proc cas;
decisionTree.forestTrain /
inputs={{name="age"},
{name="amount"},
{name="coapp"},
{name="duration"},
{name="foreign"},
{name="job"}},
maxLevel="5",
nominals={{name="coapp"}, {name="foreign"}, {name="job"}},
saveState={name="FOREST_ASTORE", replace="True"},
seed="1234",
table={name="DMAGECR"},
target="good_bad";
run;
The following code runs the score action from the astore action set to score the forest model on the dmagecr data table in order to produce the forest_scored table, while keeping all the input and ID variables:
proc cas;
astore.score /
casout={name="FOREST_SCORED", replace="True"},
copyVars={"age", "amount", "coapp", "duration", "foreign", "job", "id"},
rstore="FOREST_ASTORE",
table="DMAGECR";
run;
The following code runs the linearExplainer action to explain the prescored data that are saved in the forest_scored data table:
proc cas;
explainModel.linearExplainer /
inputs={{name="age"},
{name="amount"},
{name="coapp"},
{name="duration"},
{name="foreign"},
{name="job"}},
nominals={{name="coapp"}, {name="foreign"}, {name="job"}},
predictedTarget="P_good_badbad",
preset="KERNELSHAP",
query={name="FOREST_SCORED", where="id=20"},
seed="1234",
table={name="FOREST_SCORED"};
run;
The table parameter names the input reference data table. The query parameter names the input query data table. The predictedTarget parameter specifies that the variable P_good_badbad be used as the predicted target variable. The seed parameter specifies the seed to use for pseudorandom number generation. The preset parameter specifies that the preset explanation method is the Kernel SHAP method. The inputs parameter specifies that the variables age, amount, coapp, duration, foreign, and job be used as inputs. The nominals parameter specifies that the variables coapp, foreign, and job be used as nominal variables. Because the reference and query data tables are prescored and both contain the predicted target variable, you can run the linearExplainer action without specifying a model.
The "Explainer Information" table, "Parameter Estimates" table, and "Explainer Fidelity" table that these statements produce are shown in Output 13.5.1, Output 13.5.2, and Output 13.5.3, respectively.
Output 13.5.1: Explainer Information
| Explainer Information | |||
|---|---|---|---|
| RowId | Description | cValue | Value |
| DATAGEN | Data Generation Method | None | . |
| DISTANCE | Distance Method | SHAP Kernel | . |
| EXPLAINER | Explainer Type | Regression | . |
| BINARYENCODING | Binary Encoding | All | . |
| STANDARDIZE | Standardize Parameter Estimates | None | . |
| INCLUDEMISSING | Include Missing as a Level | No | . |
| BINWIDTH | Bin Width | 0.1 | 0.1 |
| SEED | Seed | 1234 | 1234 |
Output 13.5.2: Explainer Parameter Estimates
| Parameter Estimates | ||||||||
|---|---|---|---|---|---|---|---|---|
| Variable | BinaryVariable | coapp | foreign | job | Estimate | QueryValue | LowerBound | UpperBound |
| Intercept | . | . | . | 0.0999985731 | . | . | . | |
| age | _bCode_AGE_ | . | . | . | -0.006124607 | 31 | 29.862453143 | 32.137546857 |
| amount | _bCode_AMOUNT_ | . | . | . | -0.021163095 | 3430 | 3147.7263124 | 3712.2736876 |
| coapp | _bCode_COAPP_ | 1 | . | . | -0.061515806 | 1 | . | . |
| duration | _bCode_DURATION_ | . | . | . | -0.03917267 | 24 | 22.794118555 | 25.205881445 |
| foreign | _bCode_FOREIGN_ | . | 1 | . | 0.0535040643 | 1 | . | . |
| job | _bCode_JOB_ | . | . | 3 | -0.025525898 | 1 | . | . |
Output 13.5.3: Explainer Fidelity Information
| Explainer Fidelity | ||
|---|---|---|
| ModelPred | ExplainerPred | ExplainerRMSE |
| 0 | 5.6167123E-7 | 0.1320630939 |
Compute the Shapley Values (Kernel SHAP Method) Using Prescored Data
This section contains Lua code for the analysis in the CASL version of this example, which contains details about the results.
Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:
s:loadtable{casLib="casuser", path="dmagecr.csv"}
For more information about coding in Lua, see Getting Started with SAS Viya for Lua and SAS Viya: System Programming Guide.
This example runs the linearExplainer action to explain a prediction that is made by a forest model by using the Kernel SHAP method with prescored data.
The following code runs the forestTrain action in the decisionTree action set to build a forest model in order to predict whether the credit rating of each individual in the dmagecr data table is good or bad. The action also outputs an analytic store named forest_astore.
s:decisionTree_forestTrain{
table = {name ="DMAGECR"},
seed = 1234,
saveState = {name = 'FOREST_ASTORE',
replace = True},
target = 'good_bad',
maxLevel = 5,
inputs = {{name = "age"},
{name = "amount"},
{name = "coapp"},
{name = "duration"},
{name = "foreign"},
{name = "job"}},
nominals = {{name = "coapp"},
{name = "foreign"},
{name = "job"}}
}
The following code runs the score action from the astore action set to score the forest model on the dmagecr data table in order to produce the forest_scored table, while keeping all the input and ID variables:
s:astore.score{
table = 'DMAGECR',
rstore = 'FOREST_ASTORE',
casout = {name = 'FOREST_SCORED',
replace = True},
copyVars = {'age',
'amount',
'coapp',
'duration',
'foreign',
'job',
'id'}
}
The following code runs the linearExplainer action to explain the prescored data that are saved in the forest_scored data table:
s:explainModel_linearExplainer{
table = {name="FOREST_SCORED"},
query = {name="FOREST_SCORED", where="id=20"},
predictedTarget = "P_good_badbad",
seed = 1234,
preset = "KERNELSHAP",
inputs = {{name = "age"},
{name = "amount"},
{name = "coapp"},
{name = "duration"},
{name = "foreign"},
{name = "job"}},
nominals = {{name = "coapp"},
{name = "foreign"},
{name = "job"}}
}
The table parameter names the input reference data table. The query parameter names the input query data table. The predictedTarget parameter specifies that the variable P_good_badbad be used as the predicted target variable. The seed parameter specifies the seed to use for pseudorandom number generation. The preset parameter specifies that the preset explanation method is the Kernel SHAP method. The inputs parameter specifies that the variables age, amount, coapp, duration, foreign, and job be used as inputs. The nominals parameter specifies that the variables coapp, foreign, and job be used as nominal variables. Because the reference and query data tables are prescored and both contain the predicted target variable, you can run the linearExplainer action without specifying a model.
For example output from the action, see the SAS code examples.
Compute the Shapley Values (Kernel SHAP Method) Using Prescored Data
This section contains Python code for the analysis in the CASL version of this example, which contains details about the results.
Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:
s.upload_file('dmagecr.csv')
For more information about coding in Python, see Getting Started with SAS Viya for Python and SAS Viya: System Programming Guide.
This example runs the linearExplainer action to explain a prediction that is made by a forest model by using the Kernel SHAP method with prescored data.
The following code runs the forestTrain action in the decisionTree action set to build a forest model in order to predict whether the credit rating of each individual in the dmagecr data table is good or bad. The action also outputs an analytic store named forest_astore.
s.decisiontree.foresttrain(
inputs=[dict(name='age'),
dict(name='amount'),
dict(name='coapp'),
dict(name='duration'),
dict(name='foreign'),
dict(name='job')],
maxLevel='5',
nominals=[dict(name='coapp'), dict(name='foreign'), dict(name='job')],
saveState=dict(name='FOREST_ASTORE', replace='True'),
seed='1234',
table=dict(name='DMAGECR'),
target='good_bad')
The following code runs the score action from the astore action set to score the forest model on the dmagecr data table in order to produce the forest_scored table, while keeping all the input and ID variables:
s.astore.score(
casout=dict(name='FOREST_SCORED', replace='True'),
copyVars=['age', 'amount', 'coapp', 'duration', 'foreign', 'job', 'id'],
rstore='FOREST_ASTORE',
table='DMAGECR')
The following code runs the linearExplainer action to explain the prescored data that are saved in the forest_scored data table:
s.explainmodel.linearexplainer(
inputs=[dict(name='age'),
dict(name='amount'),
dict(name='coapp'),
dict(name='duration'),
dict(name='foreign'),
dict(name='job')],
nominals=[dict(name='coapp'), dict(name='foreign'), dict(name='job')],
predictedTarget='P_good_badbad',
preset='KERNELSHAP',
query=dict(name='FOREST_SCORED', where='id=20'),
seed='1234',
table=dict(name='FOREST_SCORED'))
The table parameter names the input reference data table. The query parameter names the input query data table. The predictedTarget parameter specifies that the variable P_good_badbad be used as the predicted target variable. The seed parameter specifies the seed to use for pseudorandom number generation. The preset parameter specifies that the preset explanation method is the Kernel SHAP method. The inputs parameter specifies that the variables age, amount, coapp, duration, foreign, and job be used as inputs. The nominals parameter specifies that the variables coapp, foreign, and job be used as nominal variables. Because the reference and query data tables are prescored and both contain the predicted target variable, you can run the linearExplainer action without specifying a model.
For example output from the action, see the SAS code examples.
Compute the Shapley Values (Kernel SHAP Method) Using Prescored Data
This section contains R code for the analysis in the CASL version of this example, which contains details about the results.
Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:
m <- cas.read.csv(s, "dmagecr.csv", casOut=list(name="dmagecr"))
For more information about coding in R, see Getting Started with SAS Viya for R and SAS Viya: System Programming Guide.
This example runs the linearExplainer action to explain a prediction that is made by a forest model by using the Kernel SHAP method with prescored data.
The following code runs the forestTrain action in the decisionTree action set to build a forest model in order to predict whether the credit rating of each individual in the dmagecr data table is good or bad. The action also outputs an analytic store named forest_astore.
cas.decisionTree.forestTrain(s,
inputs=list(list(name='age'),
list(name='amount'),
list(name='coapp'),
list(name='duration'),
list(name='foreign'),
list(name='job')),
maxLevel='5',
nominals=list(list(name='coapp'), list(name='foreign'), list(name='job')),
saveState=list(name='FOREST_ASTORE', replace='True'),
seed='1234',
table=list(name='DMAGECR'),
target='good_bad')
The following code runs the score action from the astore action set to score the forest model on the dmagecr data table in order to produce the forest_scored table, while keeping all the input and ID variables:
cas.astore.score(s,
casout=list(name='FOREST_SCORED', replace='True'),
copyVars=c('age', 'amount', 'coapp', 'duration', 'foreign', 'job', 'id'),
rstore='FOREST_ASTORE',
table='DMAGECR')
The following code runs the linearExplainer action to explain the prescored data that are saved in the forest_scored data table:
cas.explainModel.linearExplainer(s,
inputs=list(list(name='age'),
list(name='amount'),
list(name='coapp'),
list(name='duration'),
list(name='foreign'),
list(name='job')),
nominals=list(list(name='coapp'), list(name='foreign'), list(name='job')),
predictedTarget='P_good_badbad',
preset='KERNELSHAP',
query=list(name='FOREST_SCORED', where='id=20'),
seed='1234',
table=list(name='FOREST_SCORED'))
The table parameter names the input reference data table. The query parameter names the input query data table. The predictedTarget parameter specifies that the variable P_good_badbad be used as the predicted target variable. The seed parameter specifies the seed to use for pseudorandom number generation. The preset parameter specifies that the preset explanation method is the Kernel SHAP method. The inputs parameter specifies that the variables age, amount, coapp, duration, foreign, and job be used as inputs. The nominals parameter specifies that the variables coapp, foreign, and job be used as nominal variables. Because the reference and query data tables are prescored and both contain the predicted target variable, you can run the linearExplainer action without specifying a model.
For example output from the action, see the SAS code examples.