The GROW statement specifies the criterion by which to split a parent node into child nodes. As it grows the tree, PROC TREESPLIT calculates the specified criterion for each predictor variable and then splits on the predictor variable that optimizes the specified criterion.
For categorical responses, the available criteria are CHAID, CHISQUARE, ENTROPY, GINI, and IGR; the default is IGR. For continuous responses, the available criteria are CHAID, FTEST, and RSS; the default is RSS.
For either categorical or continuous responses, you can specify the following criterion:
-
CHAID <(options)>
-
for categorical predictor variables, CHAID uses the value (as specified in the ALPHA= option) of a chi-square statistic (for a classification tree) or an F statistic (for a regression tree) to merge similar levels of the predictor variable until the number of children in the proposed split reaches the number that you specify in the MAXBRANCH= option. The p-values for the final split determine the variable on which to split.
For continuous predictor variables, CHAID chooses the best single split until the number of children in the proposed split reaches the value that you specify in the MAXBRANCH= option.
You can specify the following options:
-
ALPHA=value
-
specifies the maximum p-value for a split to be considered.
By default, ALPHA=0.3.
-
BONFERRONI
-
requests a Bonferroni adjustment to the p-value for a variable after the split has been determined.
By default, no adjustment is made.
For categorical responses only, you can specify the following criteria:
-
CHISQUARE <(options)>
-
uses a chi-square statistic to split each variable and then uses the p-values that correspond to the resulting splits to determine the splitting variable.
You can specify the following options:
-
ALPHA=value
-
specifies the maximum p-value for a split to be considered.
By default, ALPHA=0.3.
-
BONFERRONI
-
requests a Bonferroni adjustment to the p-value for a variable after the split has been determined.
By default, no adjustment is made.
-
ENTROPY <option>
GAIN <option>
-
uses the gain in information (decrease in entropy) to split each variable and then to determine the split. You can specify the following option:
-
MINENTROPY=number
MINGAIN=number
specifies the minimum gain value to validate a split.
-
GINI
uses the decrease in the Gini index to split each variable and then to determine the split.
-
IGR
uses the entropy metric to split each variable and then uses the information gain ratio to determine the split.
The default criterion for categorical responses is IGR.
For continuous responses only, you can specify the following criteria:
-
FTEST <(options)>
-
uses an F statistic to split each variable and then uses the resulting p-value to determine the split variable.
You can specify the following options:
-
ALPHA=value
-
specifies the maximum p-value for a split to be considered.
By default, ALPHA=0.3.
-
BONFERRONI
-
requests a Bonferroni adjustment to the p-value for a variable after the split has been determined.
By default, no adjustment is made.
-
VARIANCE
uses the change in response variance to split each variable and then to determine the split.
The default criterion for continuous responses is RSS.
You can tune the criterion by using the AUTOTUNE statement.