GAMSELECT Procedure

Example 13.2 Options to Control Boosting Model Fit

(View the complete code for this example.)

This example and Example 13.3 demonstrate how you can build an additive model and how you can control the complexity of the model that PROC GAMSELECT fits. The following DATA step creates the data table one in the current CAS engine libref named mylib:

data mylib.one;
   call streaminit(1);
   array v{200} v1-v200;
   do i=1 to 10000;
      do j=1 to 200;
          v{j}=rand("normal");
      end;
      f1=-2*sin(2*v1);
      f2=v2*v2-1./3;
      f3=v3-0.5;
      f4=exp(-v4)+exp(-1)-1;
      linp=f1+f2+f3+f4;
      y=rand("normal",linp);
      output;
   end;
 run;

The data table one consists of 10,000 observations for a continuous response variable (y) and 200 continuous variables (v1–v200). The model for the mean of response variable y includes a sine function of the variable v1, a quadratic function of the variable v2, an exponential function of the variable v3, and a linear function of the variable v4. Output 13.2.1 shows the true function curves of the four variables that construct the response.

Output 13.2.1: True Function Curves

 True Function Curves


Without knowing the data generating process, you could form an initial analysis plan by using some basic exploratory tools. A density histogram of the response variable y (not shown) suggests that it follows a normal distribution, and a scatter matrix (not shown) indicates nonlinear relationships between some of the predictors and the response and many possible nuisance predictors. On the basis of these observations, you might decide to first use PROC REGSELECT to perform the selection. This procedure enables you to use the EFFECT statement to construct regression splines to model nonlinear terms. For more information about the procedure, see Chapter 29, REGSELECT Procedure. For more information about using the EFFECT statement in PROC REGSELECT to construct regression splines, see the section Spline Effects in Chapter 2, Shared Concepts. The following program shows how you can start your analysis by performing the stepwise selection on 200 constructed regression splines:

proc regselect data=mylib.one;
   effect spl=spline(v1-v200/separate);
   model y=spl;
   selection method=stepwise;
run;

The output from PROC REGSELECT (Output 13.2.2) shows that five effects, including the intercept, are selected and shows an average square error (ASE) of about 1.4 for the selected model.

Output 13.2.2: Selected Effects and Fit Statistics

The REGSELECT Procedure
 
Selection Details

Selected Effects:Intercept spl_v1 spl_v2 spl_v3 spl_v4

Root MSE1.18341
R-Square0.86648
Adj R-Sq0.86616
AIC13395
AICC13395
SBC3573.31162
ASE1.39697


For regression modeling of data that have nonlinear structures, it is often necessary to check partial residuals and partial predictions to determine whether the included model effects sufficiently explain the nonlinear structure. Output 13.2.3 displays the partial predictions overlaid with fitted partial prediction curves for the selected model.

In Output 13.2.3, you observe that the fitted curves explain the pattern in the partial residuals reasonably well, except for the first variable, v1. The fitted curve seems to overfit the partial residuals, especially in the tail areas. This suggests that the default regression spline construction is not sufficient for modeling the nonlinear relationship between v1 and y.

Output 13.2.3: PROC REGSELECT Fit

 PROC REGSELECT Fit


This problem has a few possible solutions. One solution is to construct multiple spline bases at different scales of knot intervals in order to approximate the nonlinear structure at finer scales and plot the partial residuals for the new models. In general, you will have to keep repeating these steps in order to obtain a model that explains the nonlinear structure well. For data that have complicated dependency structures, this might take many iterations. However, you can use PROC GAMSELECT to perform an additive model selection that uses penalized splines. Because of the spline penalty, the fitted splines are less prone to exhibit extreme boundary behaviors. This example and Example 13.3 demonstrate how you can use the procedure to perform the selection and use some options to fine-tune the selection results.

The following PROC SQL statements create a macro variable, splineTerms, which contains the list of spline terms for the variables v1–v200:

proc sql noprint;
   select cat("spline(",strip(name),")") into :splineTerms separated by ' '
   from dictionary.columns
   where libname = "MYLIB" and memname = "ONE" and
   upcase(name) like 'V%';
quit;

The following PROC GAMSELECT statements use the macro variable splineTerms to specify a model for the variable y by using the boosting method without specifying any options to control the model fit:

proc gamselect data=mylib.one plots=all;
   model y = &splineTerms;
   selection method=boosting;
run;

Output 13.2.4 shows that the boosting algorithm executed 500 iterations and selected the model that corresponds to the final boosting iteration.

Output 13.2.4: Iteration Summary for Default Model

The GAMSELECT Procedure

Iteration Summary
Iterations Executed500
Selected Iteration500
Number of Selected Effects8


The "Selected Effects" table in Output 13.2.5 shows that the four spline terms that are constructed from the variables v1–v4 are included in the final model and that at most iterations the effect that was selected for an update was one of these four terms.

Output 13.2.5: Selected Effects for Default Model

Selected Effects
EffectEntry IterationTimes Selected
Spline(v4)1187
Spline(v2)567
Spline(v1)9204
Spline(v3)1835
Spline(v15)4433
Spline(v51)4642
Spline(v148)4721
Spline(v98)4931


The model also includes four effects that do not correspond to any true signal. These effects enter the model at iterations later in the run of the boosting algorithm and are selected for a small number of iterations. Output 13.2.6 shows the fit statistics for the selected model; in this case, there is only one—the ASE.

Output 13.2.6: Fit Statistics for Default Model

Fit Statistics
Average Square Error1.03910


The smoothing component plots for the selected effects in Output 13.2.7 show well-reconstructed curves for the splines that are constructed from the variables v1–v4 and show small contributions from the remaining components.

Output 13.2.7: Smoothing Component Panel for Default Model

 Smoothing Component Panel for Default Model
External File:images/gamselex2aPlot1.png


By default, PROC GAMSELECT selects the model ModifyingAbove f With caret Superscript m Super Superscript asterisk that minimizes the ASE for the training data from the set of models that correspond to each boosting iteration. This criterion for selecting the final model might result in overfitting of the training data. To prevent overfitting of the training data, you can use the CHOOSE= option in the SELECTION statement to change the criterion that is used to select the final model. In addition to changing this criterion, you can implement early stopping of the selection process by using the STOPHORIZON= and STOPTOL= options in the SELECTION statement. For more information about how PROC GAMSELECT implements early stopping, see the section Boosting.

To demonstrate the use of early stopping, the model selection process is repeated using k-fold cross validation of the ASE as the selection criterion. PROC GAMSELECT can perform cross validation either by randomly assigning observations to folds or by using an input variable to partition the data into folds. Note that the random sampling of observations for cross validation depends on the initial random number seed and how the data are distributed across machines and threads. Therefore, even if you specify the same seed value by using the SEED= option in the PROC GAMSELECT statement, results might differ between runs of PROC GAMSELECT. The following DATA step adds a variable cvFold that can be used to partition the data into folds for cross validation:

data mylib.one / single=yes;
  set mylib.one;
  call streaminit(1848);
  cvFold = rand("table",0.2,0.2,0.2,0.2,0.2);
run;

The following PROC GAMSELECT statements perform model selection by using early stopping, k-fold cross validation to select the final model, and the same list of spline terms that is specified in the splineTerms macro variable:

proc gamselect data=mylib.one plots=all;
   model y = &splineTerms;
   selection method=boosting(choose=CV index=cvFold
             stopHorizon=10 stopTol = 0.0005);
run;

Output 13.2.8 shows that the selection process terminates early, at the 300th boosting iteration, and that the selected model includes four effects.

Output 13.2.8: Iteration Summary Using Cross Validation

The GAMSELECT Procedure

Iteration Summary
Iterations Executed300
Selected Iteration300
Number of Selected Effects4


The "Fit Statistics" table in Output 13.2.9 shows the ASE for the training data and the k-fold cross validation of the ASE. The two ASE values are similar and are larger than the ASE value for the default model (Output 13.2.6).

Output 13.2.9: Fit Statistics Using Cross Validation

Fit Statistics
Average Square Error1.07588
Cross Validation ASE1.09528


Output 13.2.10 displays the smoothing component panel for the model that is selected using cross validation and early stopping. Compared to the smoothing component panel for the default model (Output 13.2.7), the curve for Spline(v1) in Output 13.2.10 shows a worse reconstruction near the maximum and minimum values of the variable v1. This difference suggests that the fit might benefit from increasing the degrees of freedom for the penalized B-spline for Spline(v1) fit at each boosting iteration. For more information about the construction of penalized B-splines, see the section Penalized B-Splines.

Output 13.2.10: Smoothing Component Panel Using Cross Validation

 Smoothing Component Panel Using Cross Validation


The following PROC GAMSELECT statements fit a reduced model by using spline terms for the variables v1–v4 with 10 degrees of freedom for the Spline(v1) estimate fit at each boosting iteration:

proc gamselect data=mylib.one plots=all;
   model y = spline(v1 / df = 10) spline(v2) spline(v3) spline(v4);
   selection method=boosting(choose=CV index=cvFold
             stopHorizon=10 stopTol = 0.0005);
run;

Output 13.2.11 displays the "Fit Statistics" table and "Iteration Summary" table for the reduced model. The selection process terminates early, at the 219th iteration, and has an ASE closer to the ASE of the default model (Output 13.2.6).

Output 13.2.11: Iteration Summary and Fit Statistics for Reduced Model

The GAMSELECT Procedure

Iteration Summary
Iterations Executed219
Selected Iteration219
Number of Selected Effects4

Fit Statistics
Average Square Error1.04808
Cross Validation ASE1.06386


The increased degrees of freedom for Spline(v1) results in a better reconstruction than the model fit that uses the default 4 degrees of freedom with cross validation and early stopping (Output 13.2.10).

Output 13.2.12: Smoothing Component Panel for Reduced Model

 Smoothing Component Panel for Reduced Model


Last updated: June 22, 2026