The GAMSELECT Procedure
Getting Started: GAMSELECT Procedure
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
Given a response with many potential predictors, this example shows how you can use the GAMSELECT procedure not only to find which are the few predictors that really do affect the response, but also to determine what functional form that effect takes.
The following DATA step creates the data table getStarted, which consists of 20,000 observations on a continuous response variable (y) and 90 continuous variables (x1–x90), in your CAS session:
data mycas.getStarted;
drop i j w;
array x{90};
call streaminit(4321);
pi=constant("pi");
do i=1 to 20000;
u = ranuni(12345);
do j=1 to dim(x);
w = ranuni(12345);
x{j} = (w + u)/2;
end;
f1 = x1;
f2 = (2*x2-1)**2;
f3 = sin(2 * pi * x3)/(2-sin(2*pi*x3));
f4 = 0.1*sin(2*pi*x4) + 0.2*cos(2*pi*x4)+0.3*sin(2*pi*x4)**2
+ 0.4*cos(2*pi*x4)**3+0.5*sin(2*pi*x4)**3;
linp = 5*f1+3*f2+4*f3+6*f4;
y = linp + 1.32 * rand("Normal");
output;
end;
run;
These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref. For the simulation, only the variables x1–x4 affect the outcome y, and the variables x1–x90 are all positively correlated.
The following statements fit a nonparametric model by using spline transformations of the variables x1–x10:
ods graphics on;
proc gamselect data= mycas.getStarted plots=all;
model y = spline(x1 ) spline(x2 ) spline(x3 ) spline(x4 ) spline(x5 )
spline(x6 ) spline(x7 ) spline(x8 ) spline(x9 ) spline(x10)
;
selection method=boosting;
run;
The MODEL statement and the SELECTION statement are required when you use PROC GAMSELECT. The output from this analysis is presented in Figure 1 through Figure 6.
Figure 1 displays the "Model Information" table. The GAMSELECT procedure uses the boosting method to fit and select the model. For more information about the boosting selection method, see the section Boosting. The response variable y is modeled using a normal distribution whose mean is modeled by an identity link function.
Figure 1: Model Information
| Model Information | |
|---|---|
| Data Source | GETSTARTED |
| Response Variable | y |
| Distribution | Normal |
| Link Function | Identity |
| Selection Method | Boosting |
| Random Number Seed | 927865077 |
Figure 2 displays the "Number of Observations" table. All 20,000 observations in the data table are used in the analysis. For data tables that have missing or invalid values, the number of observations that are used might be less than the number of observations that are read.
Figure 2: Number of Observations
| Number of Observations Read | 20000 |
|---|---|
| Number of Observations Used | 20000 |
Figure 3 displays the "Iteration Summary" table. The boosting algorithm executed 500 iterations. The selected model corresponds to the final boosting iteration and includes four effects.
Figure 3: Iteration Summary
| Iteration Summary | |
|---|---|
| Iterations Executed | 500 |
| Selected Iteration | 500 |
| Number of Selected Effects | 4 |
Figure 4 displays the "Selected Effects" table. The table displays the effects that are included in the selected model, the first iteration at which an effect entered the model, and the number of iterations at which the effect was selected. The four spline terms that are constructed from the variables x1–x4 are selected for the final model.
Figure 4: Selected Effects
| Selected Effects | ||
|---|---|---|
| Effect | Entry Iteration | Times Selected |
| Spline(x4) | 1 | 293 |
| Spline(x3) | 9 | 113 |
| Spline(x1) | 34 | 54 |
| Spline(x2) | 36 | 40 |
Figure 5 displays the smoothing component panel for the spline terms in the selected model. It displays predicted spline curves for each of the spline terms in the selected model.
Figure 5: Smoothing Component Panel

Finally, the procedure displays the table shown in Figure 6, which shows the amount of time (in seconds) that PROC GAMSELECT required to perform the various tasks in the analysis.
Figure 6: Procedure Timing
| Task Timing | ||
|---|---|---|
| Task | Seconds | Percent |
| Parsing | 0.02 | 1.01% |
| Reading and Levelizing Data | 0.01 | 0.40% |
| Creating Basis Expansions | 0.02 | 1.54% |
| Fitting Model | 1.45 | 96.78% |
| Producing Graphics | 0.00 | 0.07% |
| Miscellaneous | 0.00 | 0.27% |
| Cleaning Up | 0.00 | 0.01% |
| Total | 1.50 | 100.00% |
The following code uses the macro function SplinePrefixList to specify a model for the variable y by using spline transformations of all 90 explanatory variables, x1–x90:
%macro SplinePrefixList(prefix,n);
%do i = 1 %to &n;
spline(&prefix.&i)
%end;
%mend;
ods graphics on;
proc gamselect data= mycas.getStarted plots=all;
model y = %SplinePrefixList(x,90);
selection method=boosting;
run;
Figure 7 and Figure 8 show that the analysis that is run using all 90 predictors selects the same model that includes only the four spline terms constructed from the variables x1–x4.
Figure 7: Selected Effects
| Selected Effects | ||
|---|---|---|
| Effect | Entry Iteration | Times Selected |
| Spline(x4) | 1 | 293 |
| Spline(x3) | 9 | 113 |
| Spline(x1) | 34 | 54 |
| Spline(x2) | 36 | 40 |
Figure 8: Smoothing Component Panel
