GAMSELECT Procedure
Example 13.2 Options to Control Boosting Model Fit
(View the complete code for this example.)
This example and Example 13.3 demonstrate how you can build an additive model and how you can control the complexity of the model that PROC GAMSELECT fits. The following DATA step creates the data table one in the current CAS engine libref named mylib:
data mylib.one;
call streaminit(1);
array v{200} v1-v200;
do i=1 to 10000;
do j=1 to 200;
v{j}=rand("normal");
end;
f1=-2*sin(2*v1);
f2=v2*v2-1./3;
f3=v3-0.5;
f4=exp(-v4)+exp(-1)-1;
linp=f1+f2+f3+f4;
y=rand("normal",linp);
output;
end;
run;
The data table one consists of 10,000 observations for a continuous response variable (y) and 200 continuous variables (v1–v200). The model for the mean of response variable y includes a sine function of the variable v1, a quadratic function of the variable v2, an exponential function of the variable v3, and a linear function of the variable v4. Output 13.2.1 shows the true function curves of the four variables that construct the response.
Output 13.2.1: True Function Curves

Without knowing the data generating process, you could form an initial analysis plan by using some basic exploratory tools. A density histogram of the response variable y (not shown) suggests that it follows a normal distribution, and a scatter matrix (not shown) indicates nonlinear relationships between some of the predictors and the response and many possible nuisance predictors. On the basis of these observations, you might decide to first use PROC REGSELECT to perform the selection. This procedure enables you to use the EFFECT statement to construct regression splines to model nonlinear terms. For more information about the procedure, see Chapter 29, REGSELECT Procedure. For more information about using the EFFECT statement in PROC REGSELECT to construct regression splines, see the section Spline Effects in Chapter 2, Shared Concepts. The following program shows how you can start your analysis by performing the stepwise selection on 200 constructed regression splines:
proc regselect data=mylib.one;
effect spl=spline(v1-v200/separate);
model y=spl;
selection method=stepwise;
run;
The output from PROC REGSELECT (Output 13.2.2) shows that five effects, including the intercept, are selected and shows an average square error (ASE) of about 1.4 for the selected model.
Output 13.2.2: Selected Effects and Fit Statistics
| Selected Effects: | Intercept spl_v1 spl_v2 spl_v3 spl_v4 |
|---|
| Root MSE | 1.18341 |
|---|---|
| R-Square | 0.86648 |
| Adj R-Sq | 0.86616 |
| AIC | 13395 |
| AICC | 13395 |
| SBC | 3573.31162 |
| ASE | 1.39697 |
For regression modeling of data that have nonlinear structures, it is often necessary to check partial residuals and partial predictions to determine whether the included model effects sufficiently explain the nonlinear structure. Output 13.2.3 displays the partial predictions overlaid with fitted partial prediction curves for the selected model.
In Output 13.2.3, you observe that the fitted curves explain the pattern in the partial residuals reasonably well, except for the first variable, v1. The fitted curve seems to overfit the partial residuals, especially in the tail areas. This suggests that the default regression spline construction is not sufficient for modeling the nonlinear relationship between v1 and y.
Output 13.2.3: PROC REGSELECT Fit

This problem has a few possible solutions. One solution is to construct multiple spline bases at different scales of knot intervals in order to approximate the nonlinear structure at finer scales and plot the partial residuals for the new models. In general, you will have to keep repeating these steps in order to obtain a model that explains the nonlinear structure well. For data that have complicated dependency structures, this might take many iterations. However, you can use PROC GAMSELECT to perform an additive model selection that uses penalized splines. Because of the spline penalty, the fitted splines are less prone to exhibit extreme boundary behaviors. This example and Example 13.3 demonstrate how you can use the procedure to perform the selection and use some options to fine-tune the selection results.
The following PROC SQL statements create a macro variable, splineTerms, which contains the list of spline terms for the variables v1–v200:
proc sql noprint;
select cat("spline(",strip(name),")") into :splineTerms separated by ' '
from dictionary.columns
where libname = "MYLIB" and memname = "ONE" and
upcase(name) like 'V%';
quit;
The following PROC GAMSELECT statements use the macro variable splineTerms to specify a model for the variable y by using the boosting method without specifying any options to control the model fit:
proc gamselect data=mylib.one plots=all;
model y = &splineTerms;
selection method=boosting;
run;
Output 13.2.4 shows that the boosting algorithm executed 500 iterations and selected the model that corresponds to the final boosting iteration.
Output 13.2.4: Iteration Summary for Default Model
| Iteration Summary | |
|---|---|
| Iterations Executed | 500 |
| Selected Iteration | 500 |
| Number of Selected Effects | 8 |
The "Selected Effects" table in Output 13.2.5 shows that the four spline terms that are constructed from the variables v1–v4 are included in the final model and that at most iterations the effect that was selected for an update was one of these four terms.
Output 13.2.5: Selected Effects for Default Model
| Selected Effects | ||
|---|---|---|
| Effect | Entry Iteration | Times Selected |
| Spline(v4) | 1 | 187 |
| Spline(v2) | 5 | 67 |
| Spline(v1) | 9 | 204 |
| Spline(v3) | 18 | 35 |
| Spline(v15) | 443 | 3 |
| Spline(v51) | 464 | 2 |
| Spline(v148) | 472 | 1 |
| Spline(v98) | 493 | 1 |
The model also includes four effects that do not correspond to any true signal. These effects enter the model at iterations later in the run of the boosting algorithm and are selected for a small number of iterations. Output 13.2.6 shows the fit statistics for the selected model; in this case, there is only one—the ASE.
Output 13.2.6: Fit Statistics for Default Model
| Fit Statistics | |
|---|---|
| Average Square Error | 1.03910 |
The smoothing component plots for the selected effects in Output 13.2.7 show well-reconstructed curves for the splines that are constructed from the variables v1–v4 and show small contributions from the remaining components.
Output 13.2.7: Smoothing Component Panel for Default Model


By default, PROC GAMSELECT selects the model that minimizes the ASE for the training data from the set of models that correspond to each boosting iteration. This criterion for selecting the final model might result in overfitting of the training data. To prevent overfitting of the training data, you can use the CHOOSE= option in the SELECTION statement to change the criterion that is used to select the final model. In addition to changing this criterion, you can implement early stopping of the selection process by using the STOPHORIZON= and STOPTOL= options in the SELECTION statement. For more information about how PROC GAMSELECT implements early stopping, see the section Boosting.
To demonstrate the use of early stopping, the model selection process is repeated using k-fold cross validation of the ASE as the selection criterion. PROC GAMSELECT can perform cross validation either by randomly assigning observations to folds or by using an input variable to partition the data into folds. Note that the random sampling of observations for cross validation depends on the initial random number seed and how the data are distributed across machines and threads. Therefore, even if you specify the same seed value by using the SEED= option in the PROC GAMSELECT statement, results might differ between runs of PROC GAMSELECT. The following DATA step adds a variable cvFold that can be used to partition the data into folds for cross validation:
data mylib.one / single=yes;
set mylib.one;
call streaminit(1848);
cvFold = rand("table",0.2,0.2,0.2,0.2,0.2);
run;
The following PROC GAMSELECT statements perform model selection by using early stopping, k-fold cross validation to select the final model, and the same list of spline terms that is specified in the splineTerms macro variable:
proc gamselect data=mylib.one plots=all;
model y = &splineTerms;
selection method=boosting(choose=CV index=cvFold
stopHorizon=10 stopTol = 0.0005);
run;
Output 13.2.8 shows that the selection process terminates early, at the 300th boosting iteration, and that the selected model includes four effects.
Output 13.2.8: Iteration Summary Using Cross Validation
| Iteration Summary | |
|---|---|
| Iterations Executed | 300 |
| Selected Iteration | 300 |
| Number of Selected Effects | 4 |
The "Fit Statistics" table in Output 13.2.9 shows the ASE for the training data and the k-fold cross validation of the ASE. The two ASE values are similar and are larger than the ASE value for the default model (Output 13.2.6).
Output 13.2.9: Fit Statistics Using Cross Validation
| Fit Statistics | |
|---|---|
| Average Square Error | 1.07588 |
| Cross Validation ASE | 1.09528 |
Output 13.2.10 displays the smoothing component panel for the model that is selected using cross validation and early stopping. Compared to the smoothing component panel for the default model (Output 13.2.7), the curve for Spline(v1) in Output 13.2.10 shows a worse reconstruction near the maximum and minimum values of the variable v1. This difference suggests that the fit might benefit from increasing the degrees of freedom for the penalized B-spline for Spline(v1) fit at each boosting iteration. For more information about the construction of penalized B-splines, see the section Penalized B-Splines.
Output 13.2.10: Smoothing Component Panel Using Cross Validation

The following PROC GAMSELECT statements fit a reduced model by using spline terms for the variables v1–v4 with 10 degrees of freedom for the Spline(v1) estimate fit at each boosting iteration:
proc gamselect data=mylib.one plots=all;
model y = spline(v1 / df = 10) spline(v2) spline(v3) spline(v4);
selection method=boosting(choose=CV index=cvFold
stopHorizon=10 stopTol = 0.0005);
run;
Output 13.2.11 displays the "Fit Statistics" table and "Iteration Summary" table for the reduced model. The selection process terminates early, at the 219th iteration, and has an ASE closer to the ASE of the default model (Output 13.2.6).
Output 13.2.11: Iteration Summary and Fit Statistics for Reduced Model
| Iteration Summary | |
|---|---|
| Iterations Executed | 219 |
| Selected Iteration | 219 |
| Number of Selected Effects | 4 |
| Fit Statistics | |
|---|---|
| Average Square Error | 1.04808 |
| Cross Validation ASE | 1.06386 |
The increased degrees of freedom for Spline(v1) results in a better reconstruction than the model fit that uses the default 4 degrees of freedom with cross validation and early stopping (Output 13.2.10).
Output 13.2.12: Smoothing Component Panel for Reduced Model
