The TREESPLIT Procedure

Getting Started: TREESPLIT Procedure

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.

This example explains basic features of the TREESPLIT procedure for building a classification tree. The data contain measurements of 13 chemical attributes for 178 samples of wine. Each wine is derived from one of three cultivars that are grown in the same area of Italy, and the goal of the analysis is a model that classifies samples into cultivar groups. The data are available from the UCI Irvine Machine Learning Repository (Lichman 2013). [3]

The following statements create a data set named Wine that contains the measurements:

data Wine;
   %let url = http://archive.ics.uci.edu/ml/machine-learning-databases;
   infile "&url/wine/wine.data" url delimiter=',';
   input Cultivar Alcohol Malic Ash Alkan Mg TotPhen
         Flav NFPhen Cyanins Color Hue ODRatio Proline;
   label Cultivar = "Cultivar"
         Alcohol  = "Alcohol"
         Malic    = "Malic Acid"
         Ash      = "Ash"
         Alkan    = "Alkalinity of Ash"
         Mg       = "Magnesium"
         TotPhen  = "Total Phenols"
         Flav     = "Flavonoids"
         NFPhen   = "Nonflavonoid Phenols"
         Cyanins  = "Proanthocyanins"
         Color    = "Color Intensity"
         Hue      = "Hue"
         ODRatio  = "OD280/OD315 of Diluted Wines"
         Proline  = "Proline";
run;

The following DATA step loads the mycas.Wine data into your CAS session. This DATA step assumes that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.

data mycas.Wine;
   set Wine;
run;

The following statements print the first 10 observations of Wine, as shown in Figure 1:

proc print data=Wine(obs=10); run;

Figure 1: Partial Listing of mycas.Wine

ObsCultivarAlcoholMalicAshAlkanMgTotPhenFlavNFPhenCyaninsColorHueODRatioProline
1114.231.712.4315.61272.803.060.282.295.641.043.921065
2113.201.782.1411.21002.652.760.261.284.381.053.401050
3113.162.362.6718.61012.803.240.302.815.681.033.171185
4114.371.952.5016.81133.853.490.242.187.800.863.451480
5113.242.592.8721.01182.802.690.391.824.321.042.93735
6114.201.762.4515.21123.273.390.341.976.751.052.851450
7114.391.872.4514.6962.502.520.301.985.251.023.581290
8114.062.152.6117.61212.602.510.311.255.051.063.581295
9114.831.642.1714.0972.802.980.291.985.201.082.851045
10113.861.352.2716.0982.983.150.221.857.221.013.551045


The variable Cultivar is a nominal categorical variable that has levels 1, 2, and 3, and the 13 attribute variables are continuous.

The following statements use the TREESPLIT procedure to create a classification tree:

ods graphics on;

proc treesplit data=mycas.Wine seed=54321;
   class Cultivar;
   model Cultivar = Alcohol Malic Ash Alkan Mg TotPhen Flav
                              NFPhen Cyanins Color Hue ODRatio Proline;
   grow entropy;
   prune costcomplexity(leaves=SE);
run;

The MODEL statement specifies Cultivar as the response variable and the variables to the right of the equal sign as the predictor variables. The inclusion of Cultivar in the CLASS statement designates it as a categorical response variable and requests a classification tree. All the predictor variables are treated as continuous variables because none are included in the CLASS statement.

The GROW and PRUNE statements control two fundamental aspects of building classification and regression trees: growing and pruning. You use the GROW statement to specify the criterion for recursively splitting parent nodes into child nodes as the tree is grown. For classification trees, the default criterion is entropy; for more information, see the section Splitting Criteria.

By default, the growth process continues until the tree reaches a maximum depth of 10 (you can specify a different limit by using the MAXDEPTH= option). The result is often a large tree that overfits the data and is likely to perform poorly in predicting future data. A recommended strategy for avoiding this problem is to prune the tree to a smaller subtree that minimizes prediction error. You can use the PRUNE statement to specify the method of pruning. The default method is cost complexity.

The default output includes several informational tables, which are shown in Figure 2 through Figure 7. The "Model Information" table in Figure 2 provides information about the model and the methods that are used to grow and prune the tree.

Figure 2: Model Information

The TREESPLIT Procedure

Model Information
Split CriterionEntropy
Pruning MethodCost Complexity
Max Branches per Node2
Max Tree Depth10
Tree Depth Before Pruning3
Tree Depth After Pruning2
Number of Leaves Before Pruning7
Number of Leaves After Pruning4


The "Observation Information" table in Figure 3 provides the numbers of observations that are read and used. These numbers are the same in this example because there are no missing values in the predictor variables and the ASSIGNMISSING= option is not set to NONE.

Figure 3: Observation Information

 Training
Number of Observations Read178
Number of Observations Used178


The plot in Figure 4 is a tool for selecting the tuning parameter for cost-complexity pruning. The parameter (indicated on the lower horizontal axis) indexes a sequence of progressively smaller subtrees that are nested within the large tree. The parameter value 0 corresponds to the fully grown tree, and positive values control the trade-off between complexity (number of leaves) and fit to the training data, as measured by average misclassification rate.

Figure 4: Misclassification Rate as a Function of Cost Complexity

 Misclassification Rate as a Function of Cost Complexity


Figure 4 shows the minimum average misclassification rate, which is obtained by 10-fold cross validation. In the plot, this value is indicated by a filled-in circle. Information about the 1-SE misclassification rate is also included because the LEAVES=SE suboption was specified.

Breiman’s 1-SE rule chooses the parameter that corresponds to the smallest subtree for which the misclassification rate is less than one standard error above the minimum misclassification rate (Breiman et al. 1984). The parameter value that corresponds to the 1-SE rule is indicated by a star. The dotted vertical line in Figure 4 indicates the chosen tree.

The tree diagram in Figure 5, which is produced by default when ODS Graphics is enabled, provides an overview of the tree as a classifier.

Figure 5: Overview Diagram of Final Tree

Overview Diagram of Final Tree


The tree is constructed by starting with all the observations in the root node (labeled 0). This node is split into two internal nodes (1 and 2, respectively). Node 1 is further split into leaf nodes (3 and 4), and node 2 is further split into leaf nodes (5 and 6).

The color of the bar for each leaf node indicates the most frequent level of Cultivar among the observations in that node; this is also the classification level that is assigned to all observations in that node. The height of the bar indicates the proportion of observations in the node that have the most frequent level. The width of the link between parent and child nodes is proportional to the number of observations in the child node.

The diagram in Figure 6 provides more detail about the nodes and splits.

Figure 6: Detailed Tree Diagram

Detailed Tree Diagram


The detailed tree diagram displays a box for each node; the box contains six lines of information, separated by a horizontal line. The proportion of each level of the predictor variable is shown below the horizontal line, and the level that has the highest proportion is also displayed above the horizontal line. Also displayed above the horizontal line are the node identifier and the number of observations that are assigned to the node.

The root node (node 0) contains 178 samples. Because the MAXBRANCH= option is not specified in the preceding statements, PROC TREESPLIT divides each node into two child nodes (MAXBRANCH=2 by default). At node 0, PROC TREESPLIT determines that the impurity of the root node is maximally decreased (as measured by the entropy criterion, which is the default) by splitting the 178 observations such that all samples for which Flav greater-than-or-equal-to 1.59 are assigned to node 2 and all samples for which Flav < 1.59 are assigned to node 1. This is the primary splitting rule for node 0. In this training phase, 63 samples are assigned to node 1 and 115 samples are assigned to node 2.

Figure 6 also indicates that node 1 contains 0 samples with level 1. The legend shows that level 1 corresponds to a Cultivar value of 1, so node 1 contains no samples of the first cultivar. Similarly, the diagram indicates that node 2 contains 0 samples with level 2, and the legend shows that level 2 corresponds to a Cultivar value of 3, so node 2 contains no samples of the third cultivar.

The primary splitting rule for node 1 consists of the variable Color and the split value 3.8032. All samples for which Color greater-than-or-equal-to 3.8032 are assigned to node 4 (which then contains 49 samples), and all samples for which Color < 3.8032 are assigned to node 3 (which then contains 14 samples).

The resulting classification tree yields simple rules for predicting the cultivar. For example, a sample for which Flav less-than 1.59 and Color < 3.8032 is predicted to be from the second cultivar (node 3 indicates that level 3 has the highest proportion of observations, and the legend shows that level 3 corresponds to a Cultivar value of 2).

Figure 6 displays the entire tree that begins with the root node and has a depth of three levels. You can use the PLOTS=ZOOMEDTREE option in the PROC TREESPLIT statement to request diagrams that begin with other nodes and have specified depths.

The table in Figure 7 displays fit statistics for the tree model.

Figure 7: Fit Statistics

The TREESPLIT Procedure

Fit Statistics for Selected
Tree
 Number
of Leaves
Misclassification
Rate
Training40.0337


The misclassification rate is the total proportion of the 178 wine samples that were misclassified. The following numbers of samples were misclassified in the terminal nodes:

  • node 3: 0

  • node 4: 1

  • node 5: 1

  • node 6: 4

So the total misclassification rate is left-parenthesis 0 plus 1 plus 1 plus 4 right-parenthesis slash 178 almost-equals 0.0337.



[3] Disclaimer: SAS may reference other websites or content or resources for use at Customer's sole discretion. SAS has no control over any websites or resources that are provided by companies or persons other than SAS. Customer acknowledges and agrees that SAS is not responsible for the availability or use of any such external sites or resources, and does not endorse any advertising, products, or other materials on or available from such websites or resources. Customer acknowledges and agrees that SAS is not liable for any loss or damage that may be incurred by Customer or its end users as a result of the availability or use of those external sites or resources, or as a result of any reliance placed by Customer or its end users on the completeness, accuracy, or existence of any advertising, products, or other materials on, or available from, such websites or resources.

Last updated: September 17, 2021