GAMSELECT Procedure
Example 13.1 Testing Model Fit with Partitioned Data
(View the complete code for this example.)
The Sashelp.JunkMail data set comes from a study that classifies whether an email is junk email (coded as 1) or not (coded as 0). The data were collected by Hewlett-Packard Labs and donated by George Forman. The data table, which is specified in the following DATA step, contains 4,601 observations, with 2 binary variables and 57 continuous explanatory variables. The response variable, Class, is a binary indicator of whether an email is considered spam or not. The partitioning variable, Test, is a binary indicator that is used to divide the data into training and testing parts. The 57 explanatory variables are continuous variables that represent frequencies of some common words and characters and lengths of uninterrupted sequences of capital letters in emails.
data mylib.JunkMail;
set Sashelp.JunkMail;
run;
In the following program, the PARTITION statement divides the data into two parts. The training data have a Test value of 0 and contain about two-thirds of the data; the rest of the data are used to evaluate the fit. The macro function SplineVList is defined to specify spline terms by using the default construction.
%macro SplineVList(vars);
%let i = 1;
%do %while (%length(%scan(&vars,&i)));
spline(%scan(&vars,&i))
%let i = %eval(&i + 1);
%end;
%mend;
proc gamselect data=mylib.JunkMail;
model Class(event='1')= %SplineVList(Make Address All _3d Our Over Remove
Internet Order Mail Receive Will People Report Addresses Free Business
Email You Credit Your Font _000 Money HP HPL George _650 Lab Labs
Telnet _857 Data _415 _85 Technology _1999 Parts PM Direct CS Meeting
Original Project RE Edu Table Conference Semicolon Paren Bracket
Exclamation Dollar Pound CapAvg CapLong CapTotal)
/ dist=binary allobs;
partition rolevar=Test(train='0' test='1');
selection method=boosting(maxIter=2000);
run;
The ALLOBS option specifies that all observations be used to construct the spline basis functions. By default, only observations that have the training data role are used to construct the spline basis functions. By default, observations with the validation or test role that have spline variable values outside the range of values that is used to construct the spline basis functions are omitted from the analysis.
The "Number of Observations" and "Response Profile" tables in Output 13.1.1 are divided into training and testing columns.
Output 13.1.1: Partitioned Counts
| Number of Observations Read | 4601 |
|---|---|
| Number of Observations Used | 4601 |
| Number of Observations Used for Training | 3065 |
| Number of Observations Used for Validation | 0 |
| Number of Observations Used for Testing | 1536 |
| Response Profile | |||||
|---|---|---|---|---|---|
| Ordered Value | Class | Total Frequency | Training Frequency | Validation Frequency | Test Frequency |
| 1 | 0 | 2788 | 1847 | 0 | 941 |
| 2 | 1 | 1813 | 1218 | 0 | 595 |
| Probability modeled is Class = 1. |
The "Iteration Summary" table in Output 13.1.2 shows that the boosting algorithm executed 2,000 iterations and selected the model that corresponds to the final iteration with 26 effects.
Output 13.1.2: Iteration Summary
| Iteration Summary | |
|---|---|
| Iterations Executed | 2000 |
| Selected Iteration | 2000 |
| Number of Selected Effects | 26 |
The fit statistics for the selected model are displayed in the "Fit Statistics" table in Output 13.1.3.
Output 13.1.3: Partitioned Fit Statistics
| Fit Statistics | |
|---|---|
| ASE (Train) | 0.05450 |
| ASE (Test) | 0.06042 |
| Misclassification Rate (Train) | 0.06623 |
| Misclassification Rate (Test) | 0.06771 |
These statistics are computed for both the training and testing data. The average square error (ASE) and the misclassification rate should be similar between the two groups when the training data are representative of the testing data; for this model, the values of these statistics seem similar between the two disjoint subsets.
The "Selected Effects" table in Output 13.1.4 shows the effects that are included in the selected model, the iteration at which the effect first entered the model, and the number of iterations at which the effect was selected for an update.
Output 13.1.4: Selected Effects
| Selected Effects | ||
|---|---|---|
| Effect | Entry Iteration | Times Selected |
| Spline(Your) | 1 | 24 |
| Spline(Dollar) | 2 | 127 |
| Spline(Exclamation) | 12 | 126 |
| Spline(Free) | 22 | 84 |
| Spline(Remove) | 25 | 113 |
| Spline(HP) | 80 | 229 |
| Spline(_000) | 81 | 80 |
| Spline(Money) | 103 | 75 |
| Spline(Our) | 105 | 89 |
| Spline(CapLong) | 134 | 75 |
| Spline(George) | 189 | 223 |
| Spline(Internet) | 218 | 64 |
| Spline(CapTotal) | 312 | 68 |
| Spline(_1999) | 323 | 127 |
| Spline(Edu) | 341 | 117 |
| Spline(Font) | 436 | 67 |
| Spline(Business) | 509 | 56 |
| Spline(Meeting) | 527 | 86 |
| Spline(RE) | 786 | 47 |
| Spline(Over) | 956 | 35 |
| Spline(Project) | 1283 | 35 |
| Spline(Address) | 1319 | 24 |
| Spline(Will) | 1618 | 14 |
| Spline(Semicolon) | 1645 | 12 |
| Spline(Conference) | 1984 | 1 |
| Spline(CapAvg) | 1986 | 2 |
For comparison, the following program uses PROC LOGSELECT to fit a parametric model by using forward selection and the same set of independent variables:
proc logselect data=mylib.JunkMail;
model Class(event='1')=Make Address All _3d Our Over Remove Internet Order
Mail Receive Will People Report Addresses Free Business Email You
Credit Your Font _000 Money HP HPL George _650 Lab Labs Telnet _857
Data _415 _85 Technology _1999 Parts PM Direct CS Meeting Original
Project RE Edu Table Conference Semicolon Paren Bracket Exclamation
Dollar Pound CapAvg CapLong CapTotal;
partition rolevar=Test(train='0' test='1');
selection method=forward;
run;
The "Selected Effects" table in Output 13.1.5 displays the 24 effects in the model that is selected by PROC LOGSELECT. The models that are selected by PROC GAMSELECT and PROC LOGSELECT have 22 variables in common.
Output 13.1.5: PROC LOGSELECT Selected Effects
| Selected Effects: | Intercept Our Over Remove Internet Order Will Free Business You Your Font _000 Money HP George Parts Meeting RE Edu Semicolon Exclamation Dollar CapAvg CapLong |
|---|
The "Fit Statistics" table for the model that PROC LOGSELECT fits is shown in Output 13.1.6. The ASE and misclassification rates for both the training and test data are smaller for the model that PROC GAMSELECT fits (Output 13.1.3).
Output 13.1.6: PROC LOGSELECT Fit Statistics
| Fit Statistics | ||
|---|---|---|
| Description | Training | Testing |
| -2 Log Likelihood | 1242.59491 | 823.68742 |
| AIC (smaller is better) | 1292.59491 | 873.68742 |
| AICC (smaller is better) | 1293.02268 | 874.54835 |
| SBC (smaller is better) | 1443.28998 | 1007.11085 |
| Average Square Error | 0.05659 | 0.06353 |
| -2 Log L (Intercept-only) | 4118.98701 | 2050.73514 |
| R-Square | 0.60877 | 0.55016 |
| Max-rescaled R-Square | 0.82359 | 0.74661 |
| McFadden's R-Square | 0.69833 | 0.59835 |
| Misclassification Rate | 0.07471 | 0.07813 |
| Difference of Means | 0.75122 | 0.73431 |