SUPERLEARNER Procedure

Getting Started: SUPERLEARNER Procedure

(View the complete code for this example.)

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.

This example demonstrates how you can use the SUPERLEARNER procedure to train a super learner model.

The Sashelp.Baseball data table contains salary and performance information for Major League Baseball players (excluding pitchers) who played at least one game in both the 1986 and 1987 seasons (Time Inc. 1987). You can load the Sashelp.Baseball data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step:

data mylib.baseball;
  set sashelp.baseball;
run;

The following statements train a super learner model:

proc superlearner data=mylib.baseball seed=24893;
   target logSalary / level=interval;
   input nAtBat nHits nHome nRuns nRBI nBB
         yrMajor crAtBat crHits crHome
         crRuns crRbi crBB nOuts nAssts nError
         / level=interval;
   input league division / level=nominal;
   baselearner 'lm' regselect;
   baselearner 'dtree' treesplit;
   baselearner 'bart' bart;
   crossvalidation kfold=6;
   output out=mylib.predout copyvar=logSalary;
run;

When you use PROC SUPERLEARNER for training, you must specify the TARGET statement, at least one INPUT statement, and at least two BASELEARNER statements.

You use the TARGET statement to specify the target variable. You use the LEVEL= option to specify that the target variable is treated as a continuous (interval) variable for the analysis. You specify two INPUT statements in this example; one statement contains continuous predictors, and the other statement contains categorical predictors. You specify whether the variables in each INPUT statement are treated as interval (continuous) or nominal (categorical) variables by using the LEVEL= option.

When you use the BASELEARNER statement to specify a base learner model, you begin by specifying a name of the base learner in single quotes. After the name, you must specify a model type. The BASELEARNER statement supports various model types that you can use to build a base learner. It also supports a large number of options that you can use to customize each base learner. For more information, see the BASELEARNER statement.

In this example, you specify three base learners by using three BASELEARNER statements. You use the REGSELECT, TREESPLIT, and BART model types to specify a linear regression model, a decision tree, and a Bayesian additive regression trees model, respectively.

You split the input data into five folds for cross-validation by specifying the KFOLD= option in the CROSSVALIDATION statement. If you omit this statement, cross-validation is still conducted, and the number of folds is computed on the basis of the size of the input data set. For more information, see the section K-Fold Cross-Validation.

To form the final predictive model, PROC SUPERLEARNER uses a meta-learning step to form a weighted combination of the predictions from the base learners. In this example, the target variable is continuous, so the meta-learning step uses the convex-constrained least squares method by default. For more information, see the section Meta-learning Methods.

Finally, you use the OUTPUT statement to create an output data table named predout that contains predicted response values from the super learner. You use the OUT= option to specify the name of the output data table. You use the COPYVAR= option to copy the target variable from the input data table to the output data table.

The output from this analysis is presented in Figure 1 and Figure 2. The "Model Information" table in Figure 1 summarizes important information about the input data and primary analysis options. Information includes the input data source, response (target) variable and its type, number of folds used in cross-validation, number of base learners, loss function, meta-learning method, and random number seed.

The second table in Figure 1 shows the number of observations in the input data table. PROC SUPERLEARNER uses only complete cases in the input data for training, so an observation is excluded from training if any analysis variable contains a missing value. For more information, see the section Missing Values.

Figure 1: Model Information and Number of Observations

Model Information
Data SourceBASEBALL
Target VariablelogSalary
Target Variable TypeINTERVAL
Number of Folds6
Number of Base Learners3
Loss FunctionQuadratic Loss
Meta-learning MethodConvex-constrained Least Squares
Random Number Seed24893

Number of Observations Read322
Number of Observations Used for Training263
Number of Observations Used for Scoring322


Figure 2 displays information about the super learner model coefficients, which describe the contribution of each base learner model in the super learner model ensemble. The cross-validated risk measures each base learner’s predictive performance in cross-validation. (For more information about cross-validated risk, see the section Cross-Validated Risk.) In this example, the "bart" base learner model has the smallest cross-validated risk and therefore shows the best out-of-sample performance. It receives the largest super learner model coefficient.

Figure 2: Super Learner Model Coefficients

Super Learner Model Coefficients
NameModel TypeCoefficientCross-Validated
Risk
lmOrdinary Linear Least Squares0.069080.35223
dtreeDecision Tree0.128850.26394
bartBayesian Additive Regression Trees0.802070.17854


The following PROC PRINT statements print the first 10 observations of the predicted and observed response values. The results are shown in Figure 3.

proc print data=mylib.predout(obs=10);
run;

Figure 3: First 10 Observations in the Output Data Table

ObsP_logSalarylogSalary
16.760286.65286
25.723265.70378
35.99384.
46.063495.76832
54.593004.70048
66.72476.
75.956495.95324
86.626326.74524
94.928264.38203
104.342544.31749


Last updated: June 22, 2026