SUPERLEARNER Procedure
Example 33.1 Storing and Scoring
(View the complete code for this example.)
When you use the SUPERLEARNER procedure to train a predictive model, you can save the model by using the STORE statement. You can then use the saved model to score new observations or compute predictive margins by fixing the values of one or more predictor variables. To do so, you use the RESTORE= option in the PROC SUPERLEARNER statement for subsequent scoring.
The following statements generate the data table inputData, which consists of 5,000 observations on a continuous response variable (y), five continuous predictor variables (x1–x5), and two binary predictor variables (z1,z2), in your CAS session:
data mylib.dataForTrain / single=yes;
drop i;
array x{5};
array z{2};
call streaminit(6520);
pi=constant("pi");
do i=1 to 5000;
x1=rand("Normal");
x2=rand("Uniform");
z1=rand("Bernoulli",0.7);
x3=rand("Normal",x1*z1,0.6*z1);
z2=rand("Bernoulli",logistic(z1*x3));
x4=rand("Normal",0,0.8*z2);
if (x4 - z1) > 0 then x5 = x4 - z1;
else x5 = (x4 - z1)**2;
y=2+x1+x1*z1+sin(pi/2*x2)+x3**2+z2+x4**3+x5+rand("Normal");
output;
end;
run;
Note that the response variable y is generated on the basis of x1–x5, z1, and z2.
The following program trains a super learner ensemble model on the simulated data:
proc superlearner data=mylib.dataForTrain seed=2324;
target y / level=interval;
input z1 z2 / level=nominal;
input x1-x5 / level=interval;
baselearner 'lm' regselect;
baselearner 'lasso_2way' regselect(selection=elasticnet(lambda=5 mixing=1))
class=(z1 z2) effect=(z1|z2|x1|x2|x3|x4|x5 @2);
baselearner 'gam' gammod class=(z1 z2) param(z1 z2 x1 x2)
spline(x3) spline(x4) spline(x5);
baselearner 'bart' bart(nTree=10 nMC=100);
baselearner 'forest' forest;
baselearner 'svm' svmachine;
baselearner 'factmac' factmac(nfactors=4 learnstep=0.15);
store out=mylib.slmodel;
run;
When you use PROC SUPERLEARNER for training, you must specify the TARGET statement, one or more INPUT statements, and two or more BASELEARNER statements.
You use the TARGET statement to specify y as the response variable. You indicate that it should be treated as a continuous variable by using the LEVEL= option. You specify two INPUT statements: the first one contains two variables, which are treated as categorical predictor variables; the second one contains five variables, which are labeled as continuous predictor variables.
When you use the BASELEARNER statement to specify a base learner model, you must first provide a name for the base learner in single quotes and then a model type. This statement supports various model types and options that you can choose for building a base learner. For more information, see the BASELEARNER statement.
In this example, you specify seven base learners by using seven BASELEARNER statements. For the first base learner, 'lm', you use the REGSELECT model type to specify a linear regression model.
You specify the second base learner, 'lasso_2way', again by using the REGSELECT model type. For this base learner, you use the SELECTION= option to specify the elastic net model selection method. You use the LAMDBA= and MIXING= suboptions to specify the lambda and mixing parameters, respectively. You use the EFFECT= option to specify that all the main effects and two-way interactions among the seven predictor variables are to be used as constructed effects for this base learner. When you specify the EFFECT= option, nominal predictor variables that you specify in the INPUT statement are no longer automatically considered as classification variables for this base learner. You include z1 and z2 in the CLASS= option so that the procedure treats them as classification variables for this base learner.
For the third base learner, you use the GAMMOD model type to specify a generalized additive model. You use the CLASS=, PARAM, and SPLINE input options to specify the model effects of the base learner. You require that z1 and z2 be treated as classification variables by including them in the CLASS= option. You include the main effects of z1, z2, x1, and x2 in the base learner by specifying the PARAM option. You include univariate spline effects of x3, x4, and x5 in the base learner by using the SPLINE option.
The remaining four base learners are defined by using the BART, FOREST, SVMACHINE, and FACTMAC model types, which correspond to specifying a Bayesian additive regression tree model, a forest model, a support vector machine model, and a factorization machine model, respectively. For the base learner 'bart', you customize the number of trees and the number of MCMC iterations by using the NTREE= and NMC= options, respectively. For the base learner 'factmac', you customize the number of factors and learning steps by using the NFACTORS= and LEARNSTEP= options, respectively.
You use the STORE statement to save the trained super learner ensemble model as an item store. You can use the saved item store to score new observations without having to retrain the model again. This is illustrated later in this example.
Output 33.1.1 displays basic information about the super learner ensemble specification, and Output 33.1.2 displays the estimated coefficient for each base learner.
Output 33.1.1: Model Information and Number of Observations
| Model Information | |
|---|---|
| Data Source | DATAFORTRAIN |
| Target Variable | y |
| Target Variable Type | INTERVAL |
| Number of Folds | 5 |
| Number of Base Learners | 7 |
| Loss Function | Quadratic Loss |
| Meta-learning Method | Convex-constrained Least Squares |
| Random Number Seed | 2324 |
| Number of Observations Read | 5000 |
|---|---|
| Number of Observations Used | 5000 |
Output 33.1.2: Super Learner Coefficients
| Super Learner Model Coefficients | |||
|---|---|---|---|
| Name | Model Type | Coefficient | Cross-Validated Risk |
| lm | Ordinary Linear Least Squares | 0 | 4.35073 |
| lasso_2way | Ordinary Linear Least Squares | 0.21575 | 1.40477 |
| gam | Generalized Additive Model | 0.71463 | 1.15176 |
| bart | Bayesian Additive Regression Trees | 0.06962 | 1.69381 |
| forest | Forest | 0 | 1.78765 |
| svm | Support Vector Machine | 0 | 4.52188 |
| factmac | Factorization Machine | 0 | 2.60129 |
By default, based on the input data sample size, the input data are divided into five folds for cross-validation. Because the target variable is continuous, by default the convex-constrained least squares meta-learning method is used to estimate the weight of each base learner in the super learner ensemble, and the quadratic loss is used to compute the cross-validated risk of each base learner.
The following DATA step code uses the same data-generating process to create new observations to be used for scoring:
data mylib.dataForScore / single=yes;
drop i;
array x{5};
array z{2};
call streaminit(2456);
pi=constant("pi");
do i=1 to 1000;
x1=rand("Normal");
x2=rand("Uniform");
z1=rand("Bernoulli",0.7);
x3=rand("Normal",x1*z1,0.6*z1);
z2=rand("Bernoulli",logistic(z1*x3));
x4=rand("Normal",0,0.8*z2);
if (x4 - z1) > 0 then x5 = x4 - z1;
else x5 = (x4 - z1)**2;
y=2+x1+x1*z1+sin(pi/2*x2)+x3**2+z2+x4**3+x5+rand("Normal");
output;
end;
run;
The following statements show how to use PROC SUPERLEARNER to score the new data by using the previously fitted model:
proc superlearner data=mylib.dataForScore restore=mylib.slmodel;
output out=mylib.scoredData1 learnerpred;
run;
You specify the fitted model in the RESTORE= option. You use the DATA= option to specify the data table of observations to score, and you use the OUTPUT statement to designate the output data table to contain the predicted responses. The LEARNERPRED option includes the predicted values from each base learner in the output data table, in addition to the predicted responses from the super learner ensemble. The RESTORE= option and OUTPUT statement are both required when you use PROC SUPERLEARNER to score data by using a previously saved model.
The following statements print the first 10 observations of the scored data table. The results are shown in Output 33.1.3.
proc print data=mylib.scoredData1(obs=10);
run;
Output 33.1.3: First 10 Observations in the Output Data Table
| Obs | P_y | lm | lasso_2way | gam | bart | forest | svm | factmac |
|---|---|---|---|---|---|---|---|---|
| 1 | 6.65683 | 6.54213 | 6.50964 | 6.71963 | 6.46835 | 6.50313 | 6.86751 | 4.30350 |
| 2 | 1.98392 | 1.85226 | 2.49392 | 1.79935 | 2.29800 | 2.18953 | 1.91218 | 2.03633 |
| 3 | 2.61899 | 3.58495 | 2.75508 | 2.57583 | 2.64024 | 2.56014 | 3.21489 | 3.15332 |
| 4 | 4.42276 | 5.64113 | 4.70199 | 4.32795 | 4.53066 | 4.32606 | 5.15384 | 4.35422 |
| 5 | 2.23673 | 2.88074 | 2.44980 | 2.22783 | 1.66777 | 2.56502 | 2.54407 | 3.13678 |
| 6 | 5.53240 | 7.17779 | 6.03486 | 5.29490 | 6.41315 | 5.92846 | 6.63275 | 6.20529 |
| 7 | 2.86278 | 3.19713 | 3.02958 | 2.86950 | 2.27690 | 2.76490 | 2.82833 | 3.28876 |
| 8 | 8.03762 | 7.80369 | 8.38578 | 7.90433 | 8.32684 | 8.12643 | 7.62153 | 7.72098 |
| 9 | 3.34112 | 4.96500 | 3.41055 | 3.36905 | 2.83928 | 3.39109 | 4.51414 | 5.07560 |
| 10 | 4.26333 | 3.58532 | 3.94204 | 4.35639 | 4.30383 | 4.30255 | 3.20928 | 3.13718 |
The P_y column contains super learner predictions, and the other columns contain predictions from each individual base learner.
You can also use the saved model to compute predictive margins. To compute a predictive margin, you use the MARGIN statement to specify a fixed value for one or more variables. Then the procedure scores each observation in an input data table that you specify in the DATA= option, but the values of the variable(s) that you specify in the MARGIN statement are replaced by the specified value(s). By default, PROC SUPERLEARNER prints the means of the predicted responses. You can also directly access the predicted response values for each set of interventions by specifying the MARGINPRED option in the OUTPUT statement to include them in the output data table.
The following code shows how to compute predictive margins and the predicted response values under interventions by using PROC SUPERLEARNER and the trained model that you saved earlier in this example:
proc superlearner data=mylib.dataForScore restore=mylib.slmodel;
margin "Scenario1" z1=0 x1=0.25;
margin "Scenario2" x1=0.25;
margin "Scenario3" x1=0.5;
margin "Scenario4" x1=0.75;
output out=mylib.scoredData2 marginpred;
run;
You specify the saved model mylib.slmodel in the RESTORE= option, and you specify the data table that contains new observations to be scored in the DATA= option. Four predictive margins are specified using the MARGIN statement. The first scenario fixes the values of z1 and x1 to 0 and 0.25, respectively. The remaining three predictive margins fix the value of x1 to 0.25, 0.5, and 0.75, respectively.
The "Scoring Information" table in Output 33.1.4 displays the names of the saved model, the data table that is used for training, the data table of scored observations, and the target variable. The "Number of Observations" table displays the number of observations that are used to score and compute the predictive margins.
Output 33.1.4: Scoring Information and Number of Observations
| Scoring Information | |
|---|---|
| Store | SLMODEL |
| Scoring Table | DATAFORSCORE |
| Training Table | DATAFORTRAIN |
| Target Variable | y |
| Number of Observations Read | 1000 |
|---|---|
| Number of Observations Used for Computing Predictive Margins | 1000 |
| Number of Observations Used for Scoring | 1000 |
The "Predictive Margins" table in Output 33.1.5 displays the means of the predicted margin estimates for the interventions that are specified by a MARGIN statement.
Output 33.1.5: Predictive Margins
| Predictive Margins | |
|---|---|
| Name | Estimate |
| Scenario1 | 4.99108 |
| Scenario2 | 5.27512 |
| Scenario3 | 5.62181 |
| Scenario4 | 6.00051 |
The following statements print the first 10 observations of the scored data table that contain predicted responses without intervention and predicted responses under each intervening scenario. The results are shown in Output 33.1.6.
proc print data=mylib.scoredData2(obs=10);
run;
Output 33.1.6: First 10 Observations in the Output Data Table
| Obs | P_y | Scenario1 | Scenario2 | Scenario3 | Scenario4 |
|---|---|---|---|---|---|
| 1 | 6.65683 | 5.55252 | 7.11563 | 7.41870 | 7.75641 |
| 2 | 1.98392 | 3.85930 | 3.85930 | 4.15335 | 4.48204 |
| 3 | 2.61899 | 3.56465 | 3.89671 | 4.24107 | 4.62008 |
| 4 | 4.42276 | 4.22765 | 4.57249 | 4.93537 | 5.33288 |
| 5 | 2.23673 | 3.57691 | 3.92740 | 4.25761 | 4.62246 |
| 6 | 5.53240 | 4.30642 | 4.33119 | 4.69125 | 5.08595 |
| 7 | 2.86278 | 4.11930 | 4.46090 | 4.77948 | 5.13271 |
| 8 | 8.03762 | 6.60072 | 7.11505 | 7.51966 | 7.94006 |
| 9 | 3.34112 | 4.40659 | 4.38477 | 4.71462 | 5.07911 |
| 10 | 4.26333 | 4.92668 | 5.26729 | 5.56103 | 5.88941 |
The P_y column again contains super learner predictions, and the other columns contain predictions from each intervening scenario.