GENSELECT Procedure

MODEL Statement

  • MODEL response <(response-options)> = <effects> </ model-options>;

  • MODEL events / trials <(response-options)> = <effects> </ model-options>;

The MODEL statement defines the statistical model in terms of a response variable (the target) or an events/trials specification, model effects that are constructed from variables in the input data table, and model-options. An intercept is included in the model by default. You can remove the intercept by specifying the NOINT option.

You can specify a single response variable that contains your response values. When you have binomial data, you can specify the events/trials form of the response, where one variable contains the number of positive responses (or events) and another variable contains the number of trials. Note that the values of both events and (trials – events) must be nonnegative and the value of trials must be positive.

For information about constructing the model effects, see the section Specification and Parameterization of Model Effects in Chapter 2, Shared Concepts.

There are two sets of options in the MODEL statement. The response-options determine how the GENSELECT procedure models probabilities for binary and multinomial data. The model-options control other aspects of model formation and inference. Table 6 summarizes these options.

Table 6: MODEL Statement Options

Option Description
Response Variable Options
DESCENDING Reverses the response categories
EVENT= Specifies the event category
ORDER= Specifies the sort order
REF= Specifies the reference category
Model Options
CENTER Centers and scales continuous main effects
CENTERLASSO Centers and scales all effects for model selection by the LASSO method and the elastic net method
CLB Requests confidence limits
DISTRIBUTION= Specifies the response distribution
INCLUDE= Includes effects in all models for model selection
INFORMATIVE Models missing values by using extra indicator variables
LINK= Specifies the link function
NOINT Suppresses the intercept
OFFSET= Specifies the offset variable
PHI= Specifies a fixed dispersion parameter
START= Includes effects in the initial model for model selection
TYPE3 Displays the Type 3 or joint tests of effects


Response Variable Options

Response variable options determine how the GENSELECT procedure models probabilities for binary and multinomial data.

You can specify the following response-options by enclosing them in parentheses after the response or trials variable.

DESCENDING
DESC

reverses the order of the response categories. If you specify both the DESCENDING and ORDER= options, PROC GENSELECT orders the response categories according to the ORDER= option and then reverses that order.

EVENT='category' | FIRST | LAST

specifies the event category for the binary and multinomial response model. PROC GENSELECT models the probability of the event category. The EVENT= option has no effect when there are more than two response categories.

You can specify one of the following:

'category'

specifies the value (formatted, if a format is applied) of the event category in quotation marks.

FIRST

designates the first ordered category as the event.

LAST

designates the last ordered category as the event.

By default, EVENT=FIRST.

For example, the following statements specify that observations whose formatted value is 1 represent events in the data. The probability that PROC GENSELECT models is thus the probability that the variable def takes the (formatted) value 1.

proc genselect data=mylib.MyData;
   class A B C;
   model def(event ='1') = A B C x1 x2 x3;
run;
ORDER=FORMATTED | FREQ | INTERNAL

specifies the sort order for the levels of the response variable. You can specify the following values:

FORMATTED

sorts the levels by external formatted value, except for numeric variables that have no explicit format, which are sorted by their unformatted (internal) value. For numeric variables for which you have supplied no explicit format (that is, for which there is no corresponding FORMAT statement in the current PROC GENSELECT run or in the DATA step that created the data table), the levels are ordered by their internal (numeric) value. The sort order is machine-dependent.

FREQ

sorts the levels by descending frequency count (levels that have the most observations come first in the order).

INTERNAL

sorts the levels by unformatted value. The sort order is machine-dependent.

By default, ORDER=FORMATTED.

For more information about sort order, see the chapter on the SORT procedure in the Base SAS Procedures Guide and the discussion of BY-group processing in SAS Programmers Guide: Essentials.

REF='category' | FIRST | LAST

specifies the reference category for the generalized logit model and the binary response model. For the generalized logit model, each logit contrasts a nonreference category with the reference category. For the binary response model, specifying one response category as the reference is the same as specifying the other response category as the event. You can specify one of the following:

'category'

specifies the value (formatted, if a format is applied) of the reference category in quotation marks.

FIRST

designates the first ordered category as the reference.

LAST

designates the last ordered category as the reference.

By default, REF=LAST.

Model Options

You can specify the following model-options after a slash (/):

CENTER

centers and scales continuous main effects for internal computation. (Continuous main effects are centered and scaled to aid in computing maximum likelihood estimates.) Parameter estimates and related statistics are always reported on the original scale. This option has no effect if you specify the LASSO model selection method or the elastic net selection method; instead, you should specify the CENTERLASSO option.

CENTERLASSO<=TRUE | FALSE>

specifies whether all effects, including categorical effects, are centered and scaled internally for the LASSO model selection method or the elastic net selection method. (Effects are centered and scaled to aid in model selection.) Parameter estimates and related statistics are always reported on the original scale. By default, CENTERLASSO=TRUE.

CLB

constructs confidence limits for each parameter estimate. The confidence level is 0.95 by default; you can change it by specifying the ALPHA= option. The CLB option is not available when you use either LASSO selection or elastic net selection.

DISTRIBUTION=keyword

specifies the response distribution for the model. The keywords and the associated distributions are shown in Table 7. For information about default and commonly used link functions for each distribution function, see Table 9.

Table 7: Built-In Distribution Functions

Keyword Distribution Function
BETA Beta
BINARY Binary
BINOMIAL Binary or binomial
EXPONENTIAL Exponential
GAMMA Gamma
GENPOISSON | GPOISSON Generalized Poisson
GEOMETRIC Geometric
IGAUSSIAN | IG Inverse Gaussian
LOGNORMAL | LOGN Lognormal
MULTINOMIAL Multinomial. For ordinal responses, it fits
a model that has a cumulative link function.
For nominal responses, it fits a model
that has a generalized logit link function.
NEGATIVEBINOMIAL | NB Negative binomial
NORMAL | GAUSSIAN Normal
POISSON Poisson
T<nu> t with nu degrees of freedom. If you
do not specify nu, a value of nu equals 3 is used.
TWEEDIE<(Tweedie-options)> Tweedie
WEIBULL Two-parameter Weibull


When DISTRIBUTION=TWEEDIE, you can specify the following Tweedie-options:

EQL

uses extended quasi-likelihood instead of Tweedie log likelihood in parameter estimation.

INITIALP=value

specifies a starting value for iterative estimation of the Tweedie power parameter.

OPTMETHOD=Tweedie-optimization-option

specifies the optimization method for iterative estimation of the Tweedie model parameters. You can specify the following Tweedie-optimization-options:

EQL

uses extended quasi-likelihood for a sample of the data, followed by extended quasi-likelihood for the full data. This is equivalent to the EQL Tweedie-option.

EQLLHOOD

uses extended quasi-likelihood for a sample of the data, followed by Tweedie log likelihood for the full data. This is the default method.

FINALLHOOD

uses a four-stage approach to estimating the Tweedie model parameters. The four stages are as follows:

  1. extended quasi-likelihood for a sample of the data

  2. Tweedie log likelihood for a sample of the data

  3. extended quasi-likelihood for the full data

  4. Tweedie log likelihood for the full data

LHOOD

uses Tweedie log likelihood be for a sample of the data, followed by Tweedie log likelihood for the full data.

P=value

specifies a value to use as a fixed Tweedie power parameter.

SAMPLEFRAC=value

specifies a value to use as the fraction of the data that are used to compute starting values for the Tweedie distribution. The value must be between 0 and 1.

INCLUDE=n
INCLUDE=single-effect
INCLUDE=(effect-list)

forces effects to be included in all models. If you specify INCLUDE=n, then the first n effects that are listed in the MODEL statement are included in all models. If you specify INCLUDE=single-effect or if you specify a list of effects within parentheses, then the specified effects are forced into all models. The effects that you specify in the INCLUDE= option must be explanatory effects that are specified in the MODEL statement before the slash (/). This option is not available if you use elastic net selection.

INFORMATIVE

models missing values by using extra model effects. These effects consist of dummy variables that take the value 1 when the value of a continuous model variable involved in the effect is missing, and take the value 0 otherwise. The missing value in the original model effect is replaced by the average value of the effect for the nonmissing values. For continuous-by-class effects, such as A*x, where A is a classification variable and x is a continuous variable, informative missingness creates multiple dummy columns and substitutes the effect mean of x that corresponds to the respective level of A. Missing values for classification variables are treated as valid levels. For more information about informative missingness, see the section Informative Missingness in Chapter 2, Shared Concepts.

LINK=keyword

specifies the link function for the model. The keywords and their associated link functions are shown in Table 8. Default and commonly used link functions for the available distributions are shown in Table 9.


For the probit and cumulative probit links, normal upper Phi Superscript negative 1 Baseline left-parenthesis dot right-parenthesis denotes the quantile function of the standard normal distribution.

If you do not specify the LINK= option, a default link function is used, as shown in Table 9. For binary or multinomial distributions, only the link functions shown in Table 9 are available. For the other distributions, you can use any link function shown in Table 8 by specifying the LINK= option. Other commonly used link functions for each distribution are shown in Table 9.


NOINT

includes no intercept in the model. An intercept is included by default. The NOINT option is not available for multinomial models.

OFFSET=variable

specifies a variable to be used as an offset to the linear predictor. An offset plays the role of an effect whose coefficient is known to be 1. The offset variable cannot appear in the CLASS statement or elsewhere in the MODEL statement. Observations that have missing values for the offset variable are excluded from the analysis.

PHI=number

specifies a fixed dispersion parameter for those distributions that have a dispersion parameter. The dispersion parameter that is used in all computations is fixed at number and not estimated.

START=n
START=single-effect
START=(effects)

begins the selection process from the designated initial model for the forward selection method. If you specify START=n, then the starting model includes the first n effects that are listed in the MODEL statement. If you specify START=single-effect or START=(effects), then the starting model includes those specified effects. The effects that you specify in the START= option must be explanatory effects that are specified in the MODEL statement before the slash (/). This option is not available when you specify METHOD=BACKWARD or ELASTICNET in the SELECTION statement.

TYPE3

computes Wald statistics for Type 3 contrasts for each effect that you specify in the MODEL statement. This option is not available for models that you fit by using the GEE method. For more information, see the section Joint Tests and Type 3 Tests.

Last updated: June 22, 2026