-
ALPHA=number
specifies a global significance level for the construction of confidence intervals. The confidence level is 1 – number. The value of number must be between 0 and 1. You can override this global significance level by specifying this option in the OUTPUT statement. By default, ALPHA=0.05.
-
APPLYROWORDER
-
uses group and order information from the input data table. You can add this information to the input table by using the partition action in the table action set. For more information, see the section The APPLYROWORDER Option in Chapter 2, Shared Concepts. By default, the APPLYROWORDER option is not enabled. If you specify this option but the table does not contain group and order information, the procedure terminates with an error.
Note: You can use group and order information for a repeated measures analysis if the group information meets certain criteria. For more information, see the section Using a Preexisting Table Order.
-
CLB
constructs confidence limits for each of the parameter estimates. The confidence level is 0.95 by default; you can change it by specifying the ALPHA= option. This option is not available when you use either LASSO selection or elastic net selection.
-
CORRB
creates the "Parameter Estimates Correlation Matrix" table. The correlation matrix is computed by normalizing the covariance matrix
. That is, if
is an element of
, then the corresponding element of the correlation matrix is
, where
. This option is not available when you use either LASSO selection or elastic net selection.
-
COVB
creates the "Parameter Estimates Covariance Matrix" table. The covariance matrix is computed as the inverse of the negative of the matrix of second derivatives of the log-likelihood function with respect to the model parameters (the Hessian matrix). This option is not available when you use either LASSO selection or elastic net selection.
-
DATA=libref.data-table
-
names the input data table for PROC GENSELECT to use. The default is the most recently created data table. libref.data-table is a two-level name, where
- libref
refers to a collection of information that is defined in the LIBNAME statement and includes the library, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about libref, see the section Using CAS Sessions and CAS Engine Librefs.
- data-table
specifies the name of the input data table.
-
FITDATA
declares the DATA= table to be the same input data table that is used for building the model, when you also specify the RESTORE= option. The FITDATA option enables you to specify the PARTFIT option to produce more fit statistics, to compute all the statistics from the OUTPUT statement that are available for your model, and to group computations according to the partition roles.
-
INPARMEST=libref.data-table
-
names an input data table that contains starting values. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
The input data table must contain an Estimate column and a ParmName column, and it can optionally contain your BY variables. You can create a version of this table by using a DISPLAYOUT statement to produce the "Parameter Estimates" table from a PROC GENSELECT program. For example:
displayout parameterestimates=pe;
You can then modify the parameter estimates in this table and read them back in by using the following statement:
proc genselect inparmest=mylib.pe;
-
ITHIST
generates the "Iteration History" table.
-
LASSORHO=r
specifies the base regularization parameter for the LASSO model selection method. The regularization parameter for step i is r
. By default, LASSORHO=0.8.
-
LASSOSTEPS=n
specifies the maximum number of steps for LASSO model selection. By default, LASSOSTEPS=20.
-
LASSOTOL=r
specifies the convergence tolerance for the optimization algorithm that solves for the LASSO parameter estimates at each step of LASSO model selection. By default, LASSOTOL=1E–6.
-
MAXRESPONSELEVELS=value
specifies the maximum number of response levels that are allowed when you have a polytomous response variable. By default, MAXRESPONSELEVELS=100.
-
NOCHECK
disables the checking process that determines whether maximum likelihood estimates of the regression parameters exist. For more information, see the section Existence of Maximum Likelihood Estimates.
-
NOCLPRINT<=number>
suppresses the display of the "Class Level Information" table if you do not specify number. If you specify number, the values of the classification variables are displayed for only those variables whose number of levels is less than number. Specifying number helps to reduce the size of the "Class Level Information" table if some classification variables have a large number of levels.
-
NOSTDERR
suppresses computation of the covariance matrix and the standard errors of the regression coefficients. When the model contains many variables (thousands), the inversion of the Hessian matrix to derive the covariance matrix and the standard errors of the regression coefficients can be time-consuming. The CORRB, COVB, and TYPE3 options are not available when the NOSTDERR option is specified. This option also disables the quasi-complete separation check; for more information, see the section Existence of Maximum Likelihood Estimates.
-
NOXPX
suppresses Hessian and
computations and invokes the NOSTDERR option. When the model contains many variables (thousands), computing the Hessian matrix during model fitting can be time-consuming. The CLB, CORRB, COVB, and TYPE3 options, the REPEATED statement, the OUTPUT statement options that rely on the covariance, and confidence limits for the parameters and for the OUTPUT and LSMEANS statements are not available when you specify the NOXPX option. The forward, backward, and stepwise selection methods are not available; however, you can specify the LASSO and elastic net methods. This option invokes the TECHNIQUE=LBFGS option by default, and the NEWRAP, NRRIDG, and TRUREG optimization techniques are not available. Because the method of identifying linearly dependent variables relies on the
matrix, some parameters that should be assigned 0 degrees of freedom can be missed.
-
PAGEOBS=number | AUTO
MAXOPTBATCH=number | AUTO
specifies the maximum number of observations to be included in a batch. During the optimization, the GENSELECT procedure reads at most number observations from the data table into memory; performs the appropriate log-likelihood, gradient, and Hessian computations on that batch of observations; then discards those observations and reads in the next batch of data for processing. Generally, a smaller number decreases memory usage but might lead to longer computation times, whereas a larger number might lead to shorter computation times but increases memory usage. The default PAGEOBS=AUTO option determines whether the entire data table can be held in a subset of your available memory; if it cannot, then number is set to 256.
-
PARTFIT
displays fit statistics in the "Fit Statistics" table that are usually produced when your data are partitioned. This option is not required when you specify a PARTITION statement. The statistic that is added to the table is the average square error (or Brier score).
-
RESTORE=libref.data-table
-
specifies the name of the analytic store that contains the context and results of a model that is selected and stored from a previous statistical analysis. The analytic store is created by a STORE statement from a previous PROC GENSELECT call or by a store parameter that is specified in a previous genmod action call.
The displayed output can include the following content:
information about the analytic store
notes that are stored by the TEXT= option of the STORE statement
a Replay section, consisting of tables that are produced when the stored model is fit
a Restore section, consisting of tables that are produced by applying the stored model to a specified input data table
You can specify the STB and CLB options in the PROC GENSELECT statement to modify the "Parameter Estimates" table; you can specify the CORRB, COVB, and TYPE3 options to display those tables; you can specify the CODE statement to produce SAS code for scoring; and you can specify LSMEANS statements to compute least squares means. If you specify a DATA= data table, then you can also specify the PARTFIT option, and you can score that data table by specifying the OUTPUT statement.
If the analytic store is created using BY-group processing, then the results are displayed according to each BY group. If you specify a DATA= input data table, then you must also specify the same BY statement that was used to create the store.
If a PARTITION statement is used to create the stored model, then the Replay section displays only information that is computed from the original training data. For the Restore section, if you create the partition for the stored model by using a ROLE variable and also specify the FITDATA option, then the ROLE variable must exist in the DATA= table and the Restore section displays results according to the roles. However, if the partitions for the stored model are created using a FRACTION option, then the roles are ignored because the roles cannot be reliably re-created; this means that output statistics that rely on the input data are not computed.
Because the selected model from a previous analysis is stored, statements and options specific to model specification, selection, and optimization are not available. In particular, the CLASS, EFFECT, FREQ, MODEL, PARTITION, REPEATED, SELECTION, STORE, and WEIGHT statements are not available to be used with the RESTORE= option.
libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
-
SEED=number
specifies an integer to be used to start the pseudorandom number generator for the LSMEANS statement’s ADJUST=SIMULATE option and for the PARTITION statement’s FRACTION option. If you do not specify a seed, or if you specify a number less than or equal to 0, the seed is generated by reading the time of day from the computer’s clock.
-
STB
-
displays the standardized estimates of the parameters in the "Parameter Estimates" table. The standardized estimate of
is given by
, where
is the total sample standard deviation for the ith explanatory variable and
The sample standard deviations for parameters that are associated with CLASS variables are computed using their codings. The standardized estimates are not computed for the intercept parameters.
-
TYPE3
computes Wald statistics for Type 3 contrasts for each effect that you specify in the MODEL statement. This option is not available for models that you fit by using the GEE method. This option is also not available when you use either LASSO selection or elastic net selection. For more information, see the section Joint Tests and Type 3 Tests.
-
USELASTITER
continues to perform computations by using the last iteration of the optimization when the optimization fails. By default, computations cease when a failure occurs.