The GAMSELECT Procedure
The OUTPUT statement creates a data table that contains observationwise statistics that PROC GAMSELECT computes after fitting the model. In order to avoid data duplication for large data tables, the variables in the input data table are not included in the output data table unless you specify them in the COPYVARS= option.
The computation of the output statistics is based on the final parameter estimates. If the model fit does not converge, missing values are produced for the quantities that depend on the estimates.
You must specify the following option:
-
OUT=CAS-libref.data-table
-
names the output data table for PROC GAMSELECT to use. You must specify this option before any other options. CAS-libref.data-table is a two-level name, where
- CAS-libref
refers to a collection of information that is defined in the LIBNAME statement and includes the caslib, which includes a path to where the data table is to be stored, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about CAS-libref, see the section Using CAS Sessions and CAS Engine Librefs.
- data-table
specifies the name of the output data table.
You can also specify the following syntax elements:
-
COMPONENT
produces componentwise statistics for all spline terms if LINP is specified as a keyword.
-
COPYVAR=variable
COPYVARS=(variables)
transfers one or more variables from the input data table to the output data table.
-
keyword <=name>
-
specifies a statistic to include in the output data table and optionally assigns a name to the variable. If you do not provide a name, PROC GAMSELECT assigns a default name based on the keyword.
You can specify the following keywords for adding statistics to the OUTPUT data table:
-
XBETA<=name>
LINP<=name>
requests the linear predictor
. For observations in which only the response variable is missing, values of the linear predictor are computed even though these observations do not affect the model fit. The default name is Xbeta.
-
PEARSON<=name>
PEARS<=name>
RESCHI<=name>
computes the Pearson residual,
, where
is the estimate of the predicted response mean and
is the response distribution variance function. The default name is Pearson.
-
PRED<=name>
PREDICTED<=name>
P<=name>
computes predicted values for the response variable. For observations in which only the response variable is missing, the predicted values are computed even though these observations do not affect the model fit. The default name is Pred.
-
RESIDUAL<=name>
RESID<=name>
R<=name>
computes the raw residual,
, where
is the estimate of the predicted mean. The default name is Residual.
-
ROLE<=name>
-
specifies the numeric variable that indicates the role played by each observation in fitting the model. The default name is _ROLE_. Table 9 shows how this variable is interpreted for each observation.
Table 9: Role Interpretation
| Value | Observation Role |
|---|
| 0 | Not used |
| 1 | Training |
| 2 | Validation |
| 3 | Testing |
If you do not partition the input data by specifying a PARTITION statement, then the role variable value is 1 for observations that are used in fitting the model and 0 for observations that have at least one missing or invalid value for the response, regressor, frequency, or weight variables.
Last updated: December 08, 2021