Shared Concepts

CLASS Statement

  • CLASS variable <(options)>ellipsis <variable <(options)>> </ global-options>;

This section applies to the following procedures: CNTSELECT, CPANEL, CQLIM, CSPATIALREG, and SEVSELECT.

The CLASS statement names the classification variables to be used as explanatory variables in the analysis. These variables enter the analysis not through their values, but through levels to which the unique values are mapped. For more information about these mappings, see the section Levelization of Classification Variables.

If the procedure permits a classification variable as a response (dependent variable or target), the response does not need to be specified in the CLASS statement.

You can specify options either as individual variable options, by enclosing the options in parentheses after the variable name, or as global-options, by placing them after a slash (/). Global-options are applied to all variables that are specified in the CLASS statement. If you specify more than one CLASS statement, the global-options that are specified in any one CLASS statement apply to all CLASS statements. However, individual CLASS variable options override the global-options.

Table 1 summarizes the values you can use for either an option or a global-option. The options are described in detail in the list that follows Table 1.

Table 1: CLASS Statement Options

Option Description
DESCENDING Reverses the sort order
MISSING Treats missing values as valid levels
ORDER= Specifies the sort order for the levels
PARAM= Specifies the parameterization of the variable
REF= Specifies the reference level of the variable
SPLIT Splits levels of CLASS variables into independent effects


DESCENDING
DESC

reverses the sort order of the classification variable. If both the DESCENDING and ORDER= options are specified, the procedure orders the categories according to the ORDER= option and then reverse that order.

MISSING

treats missing values (".", ".A", …, ".Z" for numeric variables and blanks for character variables) as valid values for the CLASS variable.

If you do not specify the MISSING option, observations that have missing values for CLASS variables are removed from the analysis.

ORDER=FORMATTED | FREQ | INTERNAL

specifies the sort order for the levels of classification variables. This ordering determines which parameters in the model correspond to each level in the data.

The following table shows how values of the ORDER= option are interpreted.

Value of ORDER= Levels Sorted By
FORMATTED External formatted values, except for numeric variables that have no explicit format, which are sorted by their unformatted (internal) values. The sort order is machine-dependent. For numeric variables for which you have supplied no explicit format, the levels are ordered by their internal values.
FREQ Descending frequency count (levels that have more observations come earlier in the order)
INTERNAL Unformatted value. The sort order is machine-dependent.

For more information about sort order, see the chapter about the SORT procedure in Base SAS Procedures Guide and the discussion of BY-group processing in the "Grouping Data" section of SAS Programmers Guide: Essentials. By default, ORDER=FORMATTED.

PARAM=keyword

specifies the parameterization method for the classification variable or variables. You can specify any of the keywords shown in the following table; design matrix columns are created from CLASS variables according to the corresponding coding schemes.

Table 2: Value of PARAM=

Value of PARAM= Coding
EFFECT Effect coding. The REF= option in the CLASS statement determines the reference level.
GLM Less-than-full-rank reference cell coding. This keyword can be used only as a global-option and is applied to all CLASS variables; all other individual variable parameterization specifications are ignored. The REF= option in the CLASS statement indirectly determines the reference level through the order of levels.
ORDINAL | 
THERMOMETER
Cumulative parameterization for an ordinal CLASS variable
POLYNOMIAL | 
POLY
Polynomial coding. If the classification variable is numeric, then the ORDER= option in the CLASS statement is ignored, and the internal unformatted values are used.
REFERENCE | 
REF
Reference cell coding. The REF= option in the CLASS statement determines the reference level.
ORTHEFFECT Orthogonalizes PARAM=EFFECT coding. The REF= option in the CLASS statement determines the reference level.
ORTHORDINAL | 
ORTHOTHERM
Orthogonalizes PARAM=ORDINAL coding
ORTHPOLY Orthogonalizes PARAM=POLYNOMIAL coding. If the classification variable is numeric, then the ORDER= option in the CLASS statement is ignored, and the internal unformatted values are used.
ORTHREF Orthogonalizes PARAM=REFERENCE coding. The REF= option in the CLASS statement determines the reference level.


All parameterizations are full rank, except for the GLM parameterization. If you specify a full rank parameterization for any CLASS variable, then every CLASS variable without a specified coding is given the EFFECT coding.

By default, PARAM=GLM. For more information about how parameterization of classification variables affects the construction and interpretation of model effects, see the section Specification and Parameterization of Model Effects.

REF='level' | keyword
REFERENCE='level' | keyword

specifies the reference level that is used when you specify a nonsingular parameterization. You can specify the following values:

'level'

specifies the level of the variable to use as the reference level. Specify the formatted value of the variable if a format is assigned. You can specify this value only for an individual variable option.

FIRST

designates the first ordered level as reference. You can specify this value either for an individual variable option or for a global-option.

LAST

designates the last ordered level as reference. You can specify this value either for an individual variable option or for a global-option.

By default, REF=LAST.

SPLIT

specifies that design matrix columns that correspond to any effect that contains a split classification variable can be selected to enter or leave a model independently of the other design columns of that effect. This option applies to procedures that perform model selection.

Suppose that the variable temp has three levels (hot, warm, and cold), that the variable gender has two levels (M and F), and that the variables are used in a PROC SEVSELECT run as follows:

proc sevselect data=mycas.data;
   loss y;
   class temp gender / split;
   scalemodel gender gender*temp;
run;

The two effects in the SCALEMODEL statement are split into eight independent effects. The effect "gender" is split into two effects that are labeled "gender_M" and "gender_F". The effect "gender*temp" is split into six effects that are labeled "gender_M*temp_hot", "gender_F*temp_hot", "gender_M*temp_warm", "gender_F*temp_warm", "gender_M*temp_cold", and "gender_F*temp_cold". The previous PROC SEVSELECT step is equivalent to the following:

proc sevselect data=mycas.data;
   loss y;
   scalemodel gender_M gender_F
             gender_M*temp_hot  gender_F*temp_hot
             gender_M*temp_warm gender_F*temp_warm
             gender_M*temp_cold gender_F*temp_cold;
run;

The SPLIT option can be used on individual classification variables. For example, consider the following PROC SEVSELECT step:

proc sevselect data=mycas.data;
   loss y;
   class temp(split) gender;
   scalemodel gender gender*temp;
run;

In this case, the effect "gender" is not split and the effect "gender*temp" is split into three effects, which are labeled "gender*temp_hot", "gender*temp_warm", and "gender*temp_cold". Furthermore, each of these three split effects now has two parameters that correspond to the two levels of "gender." The PROC SEVSELECT step is equivalent to the following:

proc sevselect data=mycas.data;
   loss y;
   class gender;
   scalemodel gender gender*temp_hot gender*temp_warm gender*temp_cold;
run;

Last updated: January 27, 2023