The GAMSELECT Procedure
SELECTION Statement
SELECTION <METHOD=method <(method-options)>>;
The SELECTION statement performs model selection by examining whether spline effects should be kept in the model.
You can specify the following methods in the SELECTION statement:
- METHOD=method <(method-options)>
specifies the method to use to select the model. You can also specify method-options that apply to the specified method by enclosing them in parentheses after the method.
The following methods are available and are explained in detail in the section Model Selection Methods. By default, METHOD=BOOSTING.
- BOOSTING
specifies the boosting method.
- SHRINKAGE
specifies the shrinkage method.
Because of the intrinsic difference between the boosting method and the shrinkage method, the two selection methods have two different sets of options that enable you to control the selection process. You can specify the following method-options for the boosting method.
- CHOOSE=CV | VALIDATE
specifies the criterion to use to select the final model. If you specify the STOPHORIZON= option, the CHOOSE= criterion is also used to evaluate early termination of the selection process. If a criterion is not specified, the average square error for the training data is used to select the final model.
You can specify the following values:
- INDEX=variable
names the variable in the input data table whose values are used to assign observations to partition folds for cross validation. This option is applicable only if the CHOOSE=CV option is specified. The number of folds equals the number of unique levels of the variable.
- KFOLD=number
specifies the number of partition folds in the random cross validation process. This option is applicable only if the CHOOSE=CV option is specified. The number must be greater than 2. By default, KFOLD=5.
- MAXITER=number
specifies the maximum number of iterations for the boosting method. By default, MAXITER=500.
-
STEPSIZE=number
LEARNINGRATE=number specifies the step size to use for the boosting algorithm, where number must be between 0 and 1. By default, STEPSIZE=0.1.
- STOPHORIZON=number
specifies the number of consecutive iterations in which the performance measured by the criterion that you specify in the CHOOSE= option must deteriorate in order to terminate the selection process. If you specify the STOPTOL= option, a relative change in the criterion is evaluated. By default, an absolute difference in this criterion is evaluated. If a criterion is not specified in the CHOOSE= option, the average square error for the training data is used.
-
STOPTOL=number
STOPTOLERANCE=number specifies the number to use as the tolerance for evaluating the relative change in the CHOOSE= option. This option is applicable only when the STOPHORIZON= option is specified. The number must be nonnegative.
You can specify the following method-options for the shrinkage method.
- ADMMABSEPS=number
specifies the absolute convergence criterion for the alternating direction method of multipliers (ADMM) algorithm that is used to compute the solution. By default, ADMMABSEPS=1E–6.
- ADMMMAXITER=number
specifies the maximum number of iterations that the ADMM algorithm can take at each step of solving a reweighted additive model. By default, ADMMMAXITER=10000.
- ADMMRELEPS=number
specifies the relative convergence criterion for the ADMM algorithm that is used to compute the solution. By default, ADMMRELEPS=1E–4.
- ADMMRHO=number
specifies the tuning parameter to use in the ADMM algorithm to control the balance between the primal residual and the dual residual. By default, number is computed from the input data.
- ADMMRHOADAPTIVE <=NO | YES>
specifies whether to use an adaptive mechanism to update the ADMM tuning parameter in the first several steps. By default, ADMMRHOADAPTIVE=YES.
- DISTRIBUTEDSEARCH <=NO | YES>
specifies whether to use distributed mode to evaluate different regularization parameter candidates in a grid environment. This mode distributes training data to workers so that each worker has a full copy of the training data and performs model fitting independently. Using this option can eliminate the cost of communication over the grid and thus improve performance. However, distributed mode requires more memory space for each worker. This option is not applicable when you are in single-machine mode. By default, DISTRIBUTEDSEARCH=YES.
-
LAMBDA1=number
L1=number sets a fixed nonnegative value to control the sparsity penalty.
-
LAMBDA2=number
L2=number sets a fixed nonnegative value to control the smoothness penalty.
-
LAMBDA3=number
L3=number sets a fixed nonnegative value to control the generalized ridge penalty.
- MAXITER=number
specifies the maximum number of iterations that generalized additive model fitting can take by solving reweighted additive models at each iteration. By default, MAXITER=500.
-
MAXLAMBDA1=number
MAXL1=number sets the maximum value to start with to search the optimal sparsity penalty parameter. By default, number is computed from the input data.
-
MAXLAMBDA2=number
MAXL2=number sets the maximum value to start with to search the optimal smoothness penalty parameter. By default, number is computed from the input data.
-
MAXLAMBDA3=number
MAXL3=number sets the maximum value to start with to search the optimal generalized ridge penalty parameter. By default, MAXLAMBDA3=0.
-
NUMLAMBDA1=number
NUML1=number sets the total number of candidates to try in searching the optimal sparsity penalty parameter. By default, NUMLAMBDA1=20.
-
NUMLAMBDA2=number
NUML2=number sets the total number of candidates to try in searching the optimal smoothness penalty parameter. By default, NUMLAMBDA2=10.
-
NUMLAMBDA3=number
NUML3=number sets the total number of candidates to try in searching the optimal generalized ridge parameter. If you specify the MAXLAMBDA3= option but not the LAMBDA3= option, then by default NUMLAMBDA3=10.
-
RHOLAMBDA1=number
RHOL1=number sets the scaling factor for the sparsity penalty parameter when computing a new candidate from its previous candidate attempt. For example, if the maximum sparsity penalty parameter is , then PROC GAMSELECT first computes the solution at , and in the following steps it computes solutions at , , and so on. By default, RHOLAMBDA1=0.8.
-
RHOLAMBDA2=number
RHOL2=number sets the scaling factor for the smoothness penalty parameter when computing a new candidate from its previous candidate attempt. For example, if the maximum smoothness penalty parameter is , then PROC GAMSELECT first computes the solution at , and in the following steps it computes solutions at , , and so on. By default, RHOLAMBDA2=0.5.
-
RHOLAMBDA3=number
RHOL3=number sets the scaling factor for the generalized ridge penalty parameter when computing a new candidate from its previous candidate attempt. For example, if the maximum generalized ridge penalty parameter is , then PROC GAMSELECT first computes the solution at , and in the following steps it computes solutions at , , and so on. If you specify the MAXLAMBDA3= option but not the LAMBDA3= option, then by default RHOLAMBDA3=0.5.