The GAMSELECT Procedure

MODEL Statement

  • MODEL response <(response-options)> = <PARAM(effects)><SPLINE(variable </ spline-options>)><SPLINE(variable1 variable2 </ spline-options>)> </ model-options>;

  • MODEL events / trials = <PARAM(effects)><SPLINE(variable </ spline-options>)><SPLINE(variable1 variable2 </ spline-options>)> </ model-options>;

The MODEL statement specifies the response (dependent or target) variable and the predictor (independent or explanatory) effects of the model. You can specify the response as a single variable or as a ratio of two variables, which are denoted as events/trials. The first form applies to all distribution families; the second form applies only to summarized binomial response data. When you have binomial data, the events variable contains the number of positive responses (or events), and the trials variable contains the number of trials. The values of both events and (trials – events) must be nonnegative, and the value of trials must be positive. If you specify a single response variable in a CLASS statement, then the response is assumed to be binary. For more information about the response-options, see the section Response Variable Options.

You can use the PARAM(effects) option to specify parametric effects that are constructed from variables in the input data; this option can appear multiple times. For information about constructing the model effects, see the section Specification and Parameterization of Model Effects in Chapter 2, Shared Concepts.

You can use the SPLINE(variable) or SPLINE(variable1 variable2) option to specify nonparametric spline effects that are constructed from variables in the input data. Only continuous variables (not classification variables) can be specified in spline effects. For more information about the options that enable you to control the construction of spline effects, see the section Spline Effect Options.

If only spline effects are present, PROC GAMSELECT builds a nonparametric additive model. If both spline and parametric effects are present, PROC GAMSELECT builds a semiparametric model by using the parametric effects as the linear part of the model. If only parametric effects are present, PROC GAMSELECT uses the boosting method to build a parametric generalized linear model by using the terms inside the parentheses of all PARAM(effects) options. The shrinkage method requires spline terms.

Three sets of options are available in the MODEL statement. The response-options determine how the GAMSELECT procedure models probabilities for binary data. The spline-options controls how each spline term forms basis expansions. The model-options control other aspects of model formation and inference. Table 4 summarizes these options, and the sections after the table describe them in detail.

Table 4: MODEL Statement Options

Option Description
Response Variable Options
DESCENDING Reverses the response categories
EVENT= Specifies the event category
ORDER= Specifies the sort order
REF= Specifies the reference category
Spline Effect Options
DEGREE= Specifies the degree of the spline effect
DETAILS Produces detailed spline information
DF= Specifies the fixed degrees of freedom
DIFFORDER= Specifies the order of the difference penalty for a B-spline
KNOTS= Specifies the knots to use for constructing the spline
LOWEREXKNOTS= Specifies the lower exterior knots
MAXKNOTS= Specifies the maximum number of knots to use for constructing the spline
SMOOTH= Specifies a fixed smoothing parameter
UPPEREXKNOTS= Specifies the upper exterior knots
WEIGHT1= Specifies the sparsity penalty strength
WEIGHT2= Specifies the smoothness penalty strength
Model Options
ALLOBS Uses values of spline variables from all data roles to construct spline basis functions
DISPERSION= Specifies the fixed dispersion parameter
DISTRIBUTION= Specifies the response distribution
INITIALPHI= Specifies the starting value of the dispersion parameter
LINK= Specifies the link function
MAXPHI= Specifies the upper bound for searching the dispersion parameter
MINPHI= Specifies the lower bound for searching the dispersion parameter
OFFSET= Specifies the offset variable


Response Variable Options

Response variable options determine how the GAMSELECT procedure models probabilities for binary data.

You can specify the following response-options by enclosing them in parentheses after the response variable:

DESCENDING
DESC

reverses the order of the response categories. If you specify both the DESCENDING and ORDER= options, PROC GAMSELECT orders the response categories according to the ORDER= option and then reverses that order.

EVENT='category' | FIRST | LAST

specifies the event category for the binary response model. PROC GAMSELECT models the probability of the event category. This option has no effect when there are more than two response categories.

You can specify any of the following values:

'category'

specifies that observations whose value matches category (formatted, if a format is applied) in quotation marks represent events in the data. For example, the following statements specify that observations that have a formatted value of '1' represent events in the data. The probability that is modeled by the GAMSELECT procedure is thus the probability that the variable def takes the (formatted) value '1'.

    proc gamselect data=mycas.MyData;
       class A B C;
       model def(event ='1') = param(A B) spline(x1) spline(x2);
       selection;
    run;
FIRST

designates the first ordered category as the event.

LAST

designates the last ordered category as the event.

By default, EVENT=FIRST.

ORDER=FORMATTED | FREQ | INTERNAL

specifies the sort order for the levels of the response variable. You can specify the following values:

FORMATTED

sorts the levels by external formatted value, except for numeric variables that have no explicit format, which are sorted by their unformatted (internal) value. For numeric variables for which you have supplied no explicit format (that is, for which there is no corresponding FORMAT statement in the current PROC GAMSELECT run or in the DATA step that created the data table), the levels are ordered by their internal (numeric) value. The sort order is machine-dependent.

FREQ

sorts the levels by descending frequency count (levels that have the most observations come first in the order).

INTERNAL

sorts the levels by unformatted value. The sort order is machine-dependent.

By default, ORDER=FORMATTED.

For more information about sort order, see the chapter about the SORT procedure in Base SAS Guide and the discussion of BY-group processing in SAS Language Reference: Concepts.

REF='category' | FIRST | LAST

specifies the reference category for the binary response model. Specifying one response category as the reference is the same as specifying the other response category as the event category. You can specify any of the following values:

'category'

specifies that observations whose value matches category (formatted, if a format is applied) are designated as the reference.

FIRST

designates the first ordered category as the reference.

LAST

designates the last ordered category as the reference.

By default, REF=LAST.

Spline Effect Options

Spline effects are specified by the SPLINE(variable / <spline-options> ) or the SPLINE(variable1 variable2 / <spline-options> ) option. Each SPLINE term defines one spline effect. The boosting selection method can take both types of specifications. The shrinkage selection method can only take SPLINE(variable / <spline-options> ). You can specify any number of SPLINE(variables) options.

Spline effect options control how each spline term forms basis expansions. For more information about the construction and estimation of spline terms, see the sections Penalized B-Splines and Natural Cubic Splines.

You can specify the following spline-options:

DEGREE=number

specifies the degree of the spline transformation, where number must be a nonnegative integer. This option is not applicable when the shrinkage method is requested. By default, DEGREE=3.

DETAILS

produces a detailed spline specification information table.

DF=number

specifies the fixed degrees of freedom for the spline at each boosting iteration. This option is applicable only when the boosting selection method is requested. The SMOOTH= option is ignored when you specify the degrees of freedom. By default, DF=4 for univariate spline terms and DF=6 for bivariate spline terms.

DIFFORDER=number

specifies the order of the difference penalty for a B-spline. This option is applicable only when the boosting selection method is requested. By default, DIFFORDER=2.

KNOTS=method

specifies the method to use for supplying user-defined knot values for constructing basis expansions. You can specify the following methods:

EQUAL(n)

specifies the number of equally spaced interior knots to use for every variable in a spline term. Boundary knots are automatically added to the knot list for each variable.

LIST(list-of-values)

specifies a list of values as knots for the spline construction. For a multivariate spline term, the listed values are taken as multiple row vectors, where each vector has values that are ordered by specified variables. If the last row vector of knots contains fewer values than the number of variables, then the last row vector is ignored. For example, the following specification of a spline term produces two actual knot vectors (bold-italic k 1 and bold-italic k 2), displayed in Table 5, and the value 5 is ignored.

spline(x1 x2/knots=list(1 2 3 4 5))

Table 5: Knot Values for a Bivariate Spline with a Supplied List

x1 x2
bold-italic k 1 1 2
bold-italic k 2 3 4


By default, the boosting method constructs basis expansions by using 20 evenly spaced interior knots for univariate spline terms, 10 evenly spaced interior knots for each variable in bivariate spline term, and evenly spaced boundary knots. By default, the shrinkage method constructs natural cubic spline basis expansions by using interior knots at unique data points and duplicated exterior knots. You can use the MAXKNOTS= option to specify the maximum number of knots if unique data points are used as interior knots.

LOWEREXKNOTS=number

specifies the lower exterior knots.

MAXKNOTS=number

specifies the maximum number of knots if data points are used to form knots. The boosting method does not support knot construction from unique data points. For the shrinkage method, if KNOTS=LIST(list-of-values) is not specified, PROC GAMSELECT forms knots from unique data points. If the number of unique data points is greater than number, a subset of size number is formed by random sampling from all unique data points. The number cannot exceed the largest integer that can be stored on the CAS server. By default, MAXKNOTS=20 if the shrinkage selection method is requested.

SMOOTH=number

specifies a fixed smoothing parameter. This option is applicable only when the boosting selection method is requested. This option is ignored when the DF= option is specified.

UPPEREXKNOTS=number

specifies the upper exterior knots.

WEIGHT1=number

specifies a sparsity penalty strength parameter. When you specify this option, a spline term might receive more or less sparsity penalty than other terms. This option is applicable only when the shrinkage selection method is requested. By default, WEIGHT1=1.

WEIGHT2=number

specifies a smoothness penalty strength parameter. When you specify this option, a spline term might receive more or less smoothness penalty than other terms. This option is applicable only when the shrinkage selection method is requested. By default, WEIGHT2=1.

Model Options

You can specify the following model-options in the MODEL statement after a slash (/):

ALLOBS

uses values of spline variables from all data roles for constructing the spline basis functions. By default, only observations that have the training data role are used to construct the spline basis functions.

DISPERSION=number
PHI=number

specifies a fixed dispersion parameter for distributions that have a dispersion parameter. The dispersion parameter that is used in all computations is fixed at number; it is not estimated.

DISTRIBUTION=keyword

specifies the response distribution for the model. The keywords and their associated distributions are shown in Table 6.

Table 6: Built-In Distribution Functions

Distribution
DISTRIBUTION= Function
BINARY | BERNOULLI Binary
BINOMIAL Binary or binomial
GAMMA Gamma
INVGAUSS | IG | IGAUSSIAN Inverse Gaussian
NORMAL | GAUSSIAN | GAUSS Normal
POISSON | POI Poisson


If you do not specify a link function in the LINK= option, a default link function is used. The default link function for each distribution is shown in Table 7. You can use any link function shown in Table 8 by specifying the LINK= option. Other commonly used link functions for each distribution are shown in Table 7.


INITIALPHI=number

specifies a starting value for iterative maximum likelihood estimation of the dispersion parameter for distributions that have a dispersion parameter.

LINK=keyword

specifies the link function for the model. The keywords and the associated link functions are shown in Table 8. Default and commonly used link functions for the available distributions are shown in Table 7.


normal upper Phi Superscript negative 1 Baseline left-parenthesis dot right-parenthesis denotes the quantile function of the standard normal distribution.

MAXPHI=number

specifies an upper bound for maximum likelihood estimation of the dispersion parameter for distributions that have a dispersion parameter.

MINPHI=number

specifies a lower bound for maximum likelihood estimation of the dispersion parameter for distributions that have a dispersion parameter.

OFFSET=variable

specifies a variable to use as an offset to the linear predictor. An offset plays the role of an effect whose coefficient is known to be 1. The offset variable cannot appear in the CLASS statement or elsewhere in the MODEL statement. Observations that have missing values for the offset variable are excluded from the analysis. For the boosting method, the offset variable is used to set the initial model estimate.

Last updated: December 08, 2021