GENSELECT Procedure

OUTPUT Statement

  • OUTPUT OUT=libref.data-table<ALL> <ALPHA=number> <COPYVARS=(variables)><keyword <=name>>…<keyword <=name>>;

The OUTPUT statement creates a data table that contains observationwise statistics that PROC GENSELECT computes after fitting the model. In order to avoid data duplication for large data tables, the variables in the input data table are not included in the output data table unless you specify them in the COPYVAR= option.

If the response variable has more than two categories, you can request the "Basic Options" listed in Table 10; the other diagnostic statistics are not available. These statistics are computed for every response category, and the automatic variable _LEVEL_ identifies the response category on which the computed values are based. That is, every observation generates several rows in the output data set. If you also specify the OBSCAT option, then the observationwise statistics are computed only for the observed response category, which is indicated by the value of the _LEVEL_ variable.

The output statistics are computed based on the final parameter estimates. If the optimization does not converge, then the output data table is not created. If you perform elastic net or LASSO selection, or if you specify the NOXPX or NOSTDERR option, then your analysis does not generate a covariance matrix, and you cannot produce the basic and diagnostic statistics in Table 10 that require a covariance matrix.

For observations in which only the response variable is missing, values of the linear predictor and the predicted values are computed even though these observations do not affect the model fit. This enables, for example, predicted values to be computed for new observations.

You must specify the following option:

OUT=libref.data-table

names the output data table for PROC GENSELECT to use. You must specify this option before any other options. libref.data-table is a two-level name, where

libref

refers to a collection of information that is defined in the LIBNAME statement and includes the library, which includes a path to where the data table is to be stored, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about libref, see the section Using CAS Sessions and CAS Engine Librefs.

data-table

specifies the name of the output data table.

You can also specify the following syntax elements:

ALL
ALLSTAT

adds all available statistics to the output data table.

ALPHA=number

specifies the significance level for the construction of confidence intervals in the output data table. The confidence level is 1 minus sans-serif-italic number. The value of number must be between 0 and 1. By default, number is equal to the value of the ALPHA= option in the PROC GENSELECT statement, or 0.05 if that option is not specified.

COPYVAR=variable
COPYVARS=(variables)

transfers one or more variables from the input data table to the output data table.

OBSCAT

requests (for multinomial models) that observationwise statistics be produced only for the observed response level. If you do not specify this option and the response variable has J levels, then the following outputs are created: for cumulative link models, J – 1 records are output for every observation in the input data that corresponds to the J – 1 lower-ordered response categories; for generalized logit models, J records are output that correspond to all J response categories.

keyword <=name>

specifies a statistic to include in the output data table and optionally names the variable name. If you do not provide a name, the GENSELECT procedure assigns a default name based on the type of statistic requested. If a statistic is not available for your model then that request is ignored.

Table 10 summarizes the keywords available in the OUTPUT statement.

Table 10: OUTPUT Statement Keywords

Keyword Description Default Name
Basic Options
INDIVIDUAL Specifies the individual predicted probabilities _IPRED_
PREDICTED Specifies the predicted probabilities _PRED_
XBETA Specifies the linear predictor _XBETA_
Basic Options Requiring the Covariance
LCL Specifies the lower confidence limit for the linear predictor _LCL_
LCLM Specifies the lower confidence limit for the event probability _LCLM_
STDXBETA Specifies the standard error estimate of the linear predictor _STDXBETA_
UCL Specifies the upper confidence limit for the linear predictor _UCL_
UCLM Specifies the upper confidence limit for the event probability _UCLM_
Diagnostic Options
RESCHI Specifies the Pearson chi-square residual _RESCHI_
RESDEV Specifies the deviance residual _RESDEV_
RESRAW Specifies the raw residual _RESRAW_
RESWORK Specifies the working residual _RESWORK_
Diagnostic Options Requiring the Covariance
CBAR Specifies the confidence interval displacement _CBAR_
DIFCHISQ Specifies the deletion chi-square goodness-of-fit change _DIFCHISQUARE_
DIFDEV Specifies the deletion deviance change _DIFDEVIANCE_
H Specifies the leverage _HATDIAG_
RESLIK Specifies the likelihood residual _RESLIK_
STDRESCHI Specifies the standardized Pearson chi-square residual _STDRESCHI_
STDRESDEV Specifies the standardized deviance residual _STDRESDEV_
Miscellaneous Options for Binary and Multinomial Response Data
INTO Names the level into which the observation is classified _INTO_
LEVEL Names the response level for a row of the output _LEVEL_
Option for Generalized Estimating Equations
CLEVERAGE Names the cluster leverage _CLEVERAGE_


The following lists describe these keywords. For more information, see the section Predicted Values and Regression Diagnostics.

CBAR

specifies the confidence interval displacement diagnostic that measures the overall change in the global regression estimates that results from deleting an individual observation. The default name is _CBAR_.

DIFCHISQ

specifies the change in the chi-square goodness-of-fit statistic that results from deleting the individual observation. The default name is _DIFCHISQUARE_.

DIFDEV

specifies the change in the deviance that results from deleting the individual observation. The default name is _DIFDEVIANCE_. This statistic is not supported for GEEs.

H

specifies the diagonal element of the hat matrix (leverage) for detecting extreme points in the design space. The default name is _HATDIAG_. For GEEs, the leverage value is the corresponding diagonal element of the cluster leverage matrix.

INDIVIDUAL
IPRED
IPROB
IP

specifies the individual predicted values for multinomial response variables. For a response variable Y with three levels, 1, 2, and 3, the individual probabilities are probability left-parenthesis sans-serif upper Y equals 1 right-parenthesis, probability left-parenthesis sans-serif upper Y equals 2 right-parenthesis, and probability left-parenthesis sans-serif upper Y equals 3 right-parenthesis. The default name is _IPRED_.

INTO<(cutpoint)>

names the variable that contains the level of the response into which an observation is classified. The default name is _INTO_. Multinomial models classify observations into the level that has the largest model-predicted probability. For binary or binomial response variables, if the predicted probability of an observation equals or exceeds the cutpoint, the observation is classified as an event; otherwise it is classified as a nonevent. You can specify the cutpoint value as a number between 0 and 1. The default value is 0.5.

LCL
LOWERXBETA

names the variable that contains the lower confidence limits for the linear predictor. The default name is _LCL_. You can set the confidence level by specifying the ALPHA= option.

LCLM
LOWERMEAN
LOWER

specifies the lower confidence limits for the mean. The default name is _LCLM_. You can set the confidence level by specifying the ALPHA= option.

LEVEL

names the variable that contains the level of the response for a given row of the output. The default name is _LEVEL_.

PREDICTED
PRED
PROB
P

specifies the predicted values for the response variable. It specifies the predicted probabilities of events for binary and nominal response variables and the cumulative predicted probabilities for ordinal response variables. For a response variable Y with three levels, 1, 2, and 3, the cumulative probabilities are probability left-parenthesis sans-serif upper Y less-than-or-equal-to 1 right-parenthesis and probability left-parenthesis sans-serif upper Y less-than-or-equal-to 2 right-parenthesis, but by default the last level, probability left-parenthesis sans-serif upper Y less-than-or-equal-to 3 right-parenthesis equals 1, is not output. The default name is _PRED_.

RESCHI
PEARSON

specifies the Pearson residual for identifying poorly fitted observations. The default name is _RESCHI_.

RESDEV

specifies the deviance residual for identifying poorly fitted observations. The default name is _RESDEV_. This statistic is not supported for GEEs.

RESLIK

specifies the likelihood residual for identifying poorly fitted observations. The default name is _RESLIK_. This statistic is not supported for GEEs.

RESRAW
RESIDUAL
R

specifies the raw residual for identifying poorly fitted observations. The default name is _RESRAW_.

RESWORK

specifies the working residual for identifying poorly fitted observations. The default name is _RESWORK_.

ROLE

specifies the numeric variable that indicates the role played by each observation in fitting the model. The default name is _ROLE_. Table 11 shows how this variable is interpreted for each observation.

Table 11: Role Interpretation

Value Observation Role
0 Not used
1 Training
2 Validation
3 Testing


If you do not partition the input data by specifying a PARTITION statement, then the role variable value is 1 for observations that are used in fitting the model and 0 for observations that have at least one missing or invalid value for the response, regressor, frequency, or weight variables.

STDRESCHI

specifies the standardized Pearson (chi-square) residual for identifying observations that are poorly accounted for by the model. The default name is _STDRESCHI_.

STDRESDEV

specifies the standardized deviance residual for identifying poorly fitted observations. The default name is _STDRESDEV_. This statistic is not supported for GEEs.

STDXBETA

specifies the standard error estimates of XBETA. The default name is _STDXBETA_.

UCL
UPPERXBETA

specifies the variable that contains the upper confidence limits for the linear predictor. The default name is _UCL_. You can set the confidence level by specifying the ALPHA= option.

UCLM
UPPERMEAN
UPPER

specifies the variable that contains the upper confidence limits for the mean. The default name is _UCLM_. You can set the confidence level by specifying the ALPHA= option.

XBETA
LINP

specifies the linear predictor. The default name is _XBETA_.

The following additional output statistic is available for GEEs. For more information about diagnostic statistics, see the section Diagnostics for Models Fit by Generalized Estimating Equations (GEEs).

CLEVERAGE
CH
CLEV

specifies the cluster leverage. The default name is _CLEVERAGE_.

Last updated: June 22, 2026