GAMSELECT Procedure

Example 13.1 Testing Model Fit with Partitioned Data

(View the complete code for this example.)

The Sashelp.JunkMail data set comes from a study that classifies whether an email is junk email (coded as 1) or not (coded as 0). The data were collected by Hewlett-Packard Labs and donated by George Forman. The data table, which is specified in the following DATA step, contains 4,601 observations, with 2 binary variables and 57 continuous explanatory variables. The response variable, Class, is a binary indicator of whether an email is considered spam or not. The partitioning variable, Test, is a binary indicator that is used to divide the data into training and testing parts. The 57 explanatory variables are continuous variables that represent frequencies of some common words and characters and lengths of uninterrupted sequences of capital letters in emails.

data mylib.JunkMail;
   set Sashelp.JunkMail;
run;

In the following program, the PARTITION statement divides the data into two parts. The training data have a Test value of 0 and contain about two-thirds of the data; the rest of the data are used to evaluate the fit. The macro function SplineVList is defined to specify spline terms by using the default construction.

%macro SplineVList(vars);
   %let i = 1;
   %do %while (%length(%scan(&vars,&i)));
      spline(%scan(&vars,&i))
      %let i = %eval(&i + 1);
   %end;
%mend;

proc gamselect data=mylib.JunkMail;
   model Class(event='1')= %SplineVList(Make Address All _3d Our Over Remove
         Internet Order Mail Receive Will People Report Addresses Free Business
         Email You Credit Your Font _000 Money HP HPL George _650 Lab Labs
         Telnet _857 Data _415 _85 Technology _1999 Parts PM Direct CS Meeting
         Original Project RE Edu Table Conference Semicolon Paren Bracket
         Exclamation Dollar Pound CapAvg CapLong CapTotal)
   / dist=binary allobs;
   partition rolevar=Test(train='0' test='1');
   selection method=boosting(maxIter=2000);
run;

The ALLOBS option specifies that all observations be used to construct the spline basis functions. By default, only observations that have the training data role are used to construct the spline basis functions. By default, observations with the validation or test role that have spline variable values outside the range of values that is used to construct the spline basis functions are omitted from the analysis.

The "Number of Observations" and "Response Profile" tables in Output 13.1.1 are divided into training and testing columns.

Output 13.1.1: Partitioned Counts

The GAMSELECT Procedure

Number of Observations Read4601
Number of Observations Used4601
Number of Observations Used for Training3065
Number of Observations Used for Validation0
Number of Observations Used for Testing1536

Response Profile
Ordered
Value
ClassTotal
Frequency
Training
Frequency
Validation
Frequency
Test
Frequency
10278818470941
21181312180595

Probability modeled is Class = 1.



The "Iteration Summary" table in Output 13.1.2 shows that the boosting algorithm executed 2,000 iterations and selected the model that corresponds to the final iteration with 26 effects.

Output 13.1.2: Iteration Summary

Iteration Summary
Iterations Executed2000
Selected Iteration2000
Number of Selected Effects26


The fit statistics for the selected model are displayed in the "Fit Statistics" table in Output 13.1.3.

Output 13.1.3: Partitioned Fit Statistics

Fit Statistics
ASE (Train)0.05450
ASE (Test)0.06042
Misclassification Rate (Train)0.06623
Misclassification Rate (Test)0.06771


These statistics are computed for both the training and testing data. The average square error (ASE) and the misclassification rate should be similar between the two groups when the training data are representative of the testing data; for this model, the values of these statistics seem similar between the two disjoint subsets.

The "Selected Effects" table in Output 13.1.4 shows the effects that are included in the selected model, the iteration at which the effect first entered the model, and the number of iterations at which the effect was selected for an update.

Output 13.1.4: Selected Effects

Selected Effects
EffectEntry IterationTimes Selected
Spline(Your)124
Spline(Dollar)2127
Spline(Exclamation)12126
Spline(Free)2284
Spline(Remove)25113
Spline(HP)80229
Spline(_000)8180
Spline(Money)10375
Spline(Our)10589
Spline(CapLong)13475
Spline(George)189223
Spline(Internet)21864
Spline(CapTotal)31268
Spline(_1999)323127
Spline(Edu)341117
Spline(Font)43667
Spline(Business)50956
Spline(Meeting)52786
Spline(RE)78647
Spline(Over)95635
Spline(Project)128335
Spline(Address)131924
Spline(Will)161814
Spline(Semicolon)164512
Spline(Conference)19841
Spline(CapAvg)19862


For comparison, the following program uses PROC LOGSELECT to fit a parametric model by using forward selection and the same set of independent variables:

proc logselect data=mylib.JunkMail;
   model Class(event='1')=Make Address All _3d Our Over Remove Internet Order
         Mail Receive Will People Report Addresses Free Business Email You
         Credit Your Font _000 Money HP HPL George _650 Lab Labs Telnet _857
         Data _415 _85 Technology _1999 Parts PM Direct CS Meeting Original
         Project RE Edu Table Conference Semicolon Paren Bracket Exclamation
         Dollar Pound CapAvg CapLong CapTotal;
   partition rolevar=Test(train='0' test='1');
   selection method=forward;
run;

The "Selected Effects" table in Output 13.1.5 displays the 24 effects in the model that is selected by PROC LOGSELECT. The models that are selected by PROC GAMSELECT and PROC LOGSELECT have 22 variables in common.

Output 13.1.5: PROC LOGSELECT Selected Effects

The LOGSELECT Procedure
 
Selection Details

Selected Effects:Intercept Our Over Remove Internet Order Will Free Business You Your Font _000 Money HP George Parts Meeting RE Edu Semicolon Exclamation Dollar CapAvg CapLong


The "Fit Statistics" table for the model that PROC LOGSELECT fits is shown in Output 13.1.6. The ASE and misclassification rates for both the training and test data are smaller for the model that PROC GAMSELECT fits (Output 13.1.3).

Output 13.1.6: PROC LOGSELECT Fit Statistics

Fit Statistics
DescriptionTrainingTesting
-2 Log Likelihood1242.59491823.68742
AIC (smaller is better)1292.59491873.68742
AICC (smaller is better)1293.02268874.54835
SBC (smaller is better)1443.289981007.11085
Average Square Error0.056590.06353
-2 Log L (Intercept-only)4118.987012050.73514
R-Square0.608770.55016
Max-rescaled R-Square0.823590.74661
McFadden's R-Square0.698330.59835
Misclassification Rate0.074710.07813
Difference of Means0.751220.73431


Last updated: June 22, 2026