TREESPLIT Procedure
Getting Started: TREESPLIT Procedure
(View the complete code for this example.)
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
This example explains basic features of the TREESPLIT procedure for building a classification tree. The data contain measurements of 13 chemical attributes for 178 samples of wine. Each wine is derived from one of three cultivars that are grown in the same area of Italy, and the goal of the analysis is a model that classifies samples into cultivar groups. The data are available from the UCI Irvine Machine Learning Repository (Lichman 2013). [5]
The following statements create a data set named Wine that contains the measurements:
data Wine;
%let url = http://archive.ics.uci.edu/ml/machine-learning-databases;
infile "&url/wine/wine.data" url delimiter=',';
input Cultivar Alcohol Malic Ash Alkan Mg TotPhen
Flav NFPhen Cyanins Color Hue ODRatio Proline;
label Cultivar = "Cultivar"
Alcohol = "Alcohol"
Malic = "Malic Acid"
Ash = "Ash"
Alkan = "Alkalinity of Ash"
Mg = "Magnesium"
TotPhen = "Total Phenols"
Flav = "Flavonoids"
NFPhen = "Nonflavonoid Phenols"
Cyanins = "Proanthocyanins"
Color = "Color Intensity"
Hue = "Hue"
ODRatio = "OD280/OD315 of Diluted Wines"
Proline = "Proline";
run;
The following DATA step loads the mylib.Wine data into your CAS session. This DATA step assumes that your CAS engine libref is named mylib, but you can substitute any appropriately defined CAS engine libref.
data mylib.Wine;
set Wine;
run;
The following statements print the first 10 observations of Wine, as shown in Figure 1:
proc print data=Wine(obs=10); run;
Figure 1: Partial Listing of mylib.Wine
| Obs | Cultivar | Alcohol | Malic | Ash | Alkan | Mg | TotPhen | Flav | NFPhen | Cyanins | Color | Hue | ODRatio | Proline |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1 | 14.23 | 1.71 | 2.43 | 15.6 | 127 | 2.80 | 3.06 | 0.28 | 2.29 | 5.64 | 1.04 | 3.92 | 1065 |
| 2 | 1 | 13.20 | 1.78 | 2.14 | 11.2 | 100 | 2.65 | 2.76 | 0.26 | 1.28 | 4.38 | 1.05 | 3.40 | 1050 |
| 3 | 1 | 13.16 | 2.36 | 2.67 | 18.6 | 101 | 2.80 | 3.24 | 0.30 | 2.81 | 5.68 | 1.03 | 3.17 | 1185 |
| 4 | 1 | 14.37 | 1.95 | 2.50 | 16.8 | 113 | 3.85 | 3.49 | 0.24 | 2.18 | 7.80 | 0.86 | 3.45 | 1480 |
| 5 | 1 | 13.24 | 2.59 | 2.87 | 21.0 | 118 | 2.80 | 2.69 | 0.39 | 1.82 | 4.32 | 1.04 | 2.93 | 735 |
| 6 | 1 | 14.20 | 1.76 | 2.45 | 15.2 | 112 | 3.27 | 3.39 | 0.34 | 1.97 | 6.75 | 1.05 | 2.85 | 1450 |
| 7 | 1 | 14.39 | 1.87 | 2.45 | 14.6 | 96 | 2.50 | 2.52 | 0.30 | 1.98 | 5.25 | 1.02 | 3.58 | 1290 |
| 8 | 1 | 14.06 | 2.15 | 2.61 | 17.6 | 121 | 2.60 | 2.51 | 0.31 | 1.25 | 5.05 | 1.06 | 3.58 | 1295 |
| 9 | 1 | 14.83 | 1.64 | 2.17 | 14.0 | 97 | 2.80 | 2.98 | 0.29 | 1.98 | 5.20 | 1.08 | 2.85 | 1045 |
| 10 | 1 | 13.86 | 1.35 | 2.27 | 16.0 | 98 | 2.98 | 3.15 | 0.22 | 1.85 | 7.22 | 1.01 | 3.55 | 1045 |
The variable Cultivar is a nominal categorical variable that has levels 1, 2, and 3, and the 13 attribute variables are continuous.
The following statements use the TREESPLIT procedure to create a classification tree:
ods graphics on;
proc treesplit data=mylib.Wine seed=54321;
class Cultivar;
model Cultivar = Alcohol Malic Ash Alkan Mg TotPhen Flav
NFPhen Cyanins Color Hue ODRatio Proline;
grow entropy;
prune costcomplexity(leaves=SE);
run;
The MODEL statement specifies Cultivar as the response variable and the variables to the right of the equals sign as the predictor variables. The inclusion of Cultivar in the CLASS statement designates it as a categorical response variable and requests a classification tree. All the predictor variables are treated as continuous variables because none are included in the CLASS statement.
The GROW and PRUNE statements control two fundamental aspects of building classification and regression trees: growing and pruning. You use the GROW statement to specify the criterion for recursively splitting parent nodes into child nodes as the tree is grown. For classification trees, the default criterion is entropy; for more information, see the section Splitting Criteria.
By default, the growth process continues until the tree reaches a maximum depth of 10 (you can specify a different limit by using the MAXDEPTH= option). The result is often a large tree that overfits the data and is likely to perform poorly in predicting future data. A recommended strategy for avoiding this problem is to prune the tree to a smaller subtree that minimizes prediction error. You can use the PRUNE statement to specify the method of pruning. The default method is cost complexity.
The default output includes several informational tables, which are shown in Figure 2 through Figure 7. The "Model Information" table in Figure 2 provides information about the model and the methods that are used to grow and prune the tree.
Figure 2: Model Information
| Model Information | |
|---|---|
| Split Criterion | Entropy |
| Pruning Method | Cost Complexity |
| Max Branches per Node | 2 |
| Max Tree Depth | 10 |
| Tree Depth Before Pruning | 3 |
| Tree Depth After Pruning | 2 |
| Number of Leaves Before Pruning | 7 |
| Number of Leaves After Pruning | 3 |
The "Observation Information" table in Figure 3 provides the numbers of observations that are read and used. These numbers are the same in this example because there are no missing values in the predictor variables and the ASSIGNMISSING= option is not set to NONE.
Figure 3: Observation Information
| Training | |
|---|---|
| Number of Observations Read | 178 |
| Number of Observations Used | 178 |
The plot in Figure 4 is a tool for selecting the tuning parameter for cost-complexity pruning. The parameter (indicated on the lower horizontal axis) indexes a sequence of progressively smaller subtrees that are nested within the large tree. The parameter value 0 corresponds to the fully grown tree, and positive values control the trade-off between complexity (number of leaves) and fit to the training data, as measured by average misclassification rate.
Figure 4: Misclassification Rate as a Function of Cost Complexity

Figure 4 shows the minimum average misclassification rate, which is obtained by 10-fold cross validation. In the plot, this value is indicated by a filled-in circle. Information about the 1-SE misclassification rate is also included because the LEAVES=SE suboption was specified.
Breiman’s 1-SE rule chooses the parameter that corresponds to the smallest subtree for which the misclassification rate is less than one standard error above the minimum misclassification rate (Breiman et al. 1984). The parameter value that corresponds to the 1-SE rule is indicated by a star. The dotted vertical line in Figure 4 indicates the chosen tree.
The tree diagram in Figure 5, which is produced by default when ODS Graphics is enabled, provides an overview of the tree as a classifier.
Figure 5: Overview Diagram of Final Tree

The tree is constructed by starting with all the observations in the root node (labeled 0). This node is split into two internal nodes (1 and 2, respectively). Node 1 is further split into leaf nodes (3 and 4), and node 2 is further split into leaf nodes (5 and 6).
The color of the bar for each leaf node indicates the most frequent level of Cultivar among the observations in that node; this is also the classification level that is assigned to all observations in that node. The height of the bar indicates the proportion of observations in the node that have the most frequent level. The width of the link between parent and child nodes is proportional to the number of observations in the child node.
The diagram in Figure 6 provides more detail about the nodes and splits.
Figure 6: Detailed Tree Diagram

The detailed tree diagram displays a box for each node; the box contains six lines of information, separated by a horizontal line. The proportion of each level of the predictor variable is shown below the horizontal line, and the level that has the highest proportion is also displayed above the horizontal line. Also displayed above the horizontal line are the node identifier and the number of observations that are assigned to the node.
The root node (node 0) contains 178 samples. Because the MAXBRANCH= option is not specified in the preceding statements, PROC TREESPLIT divides each node into two child nodes (MAXBRANCH=2 by default). At node 0, PROC TREESPLIT determines that the impurity of the root node is maximally decreased (as measured by the entropy criterion, which is the default) by splitting the 178 observations such that all samples for which Flav 1.59 are assigned to node 2 and all samples for which
Flav < 1.59 are assigned to node 1. This is the splitting rule for node 0. In this training phase, 63 samples are assigned to node 1 and 115 samples are assigned to node 2.
Figure 6 also indicates that node 1 contains 0 samples with level 1. The legend shows that level 1 corresponds to a Cultivar value of 1, so node 1 contains no samples of the first cultivar. Similarly, the diagram indicates that node 2 contains 0 samples with level 2, and the legend shows that level 2 corresponds to a Cultivar value of 3, so node 2 contains no samples of the third cultivar.
The splitting rule for node 1 consists of the variable Color and the split value 3.8032. All samples for which Color are assigned to node 4 (which then contains 49 samples), and all samples for which
Color < 3.8032 are assigned to node 3 (which then contains 14 samples).
The resulting classification tree yields simple rules for predicting the cultivar. For example, a sample for which Flav and
Color < 3.8032 is predicted to be from the second cultivar (node 3 indicates that level 3 has the highest proportion of observations, and the legend shows that level 3 corresponds to a Cultivar value of 2).
Figure 6 displays the entire tree that begins with the root node and has a depth of three levels. You can use the PLOTS=ZOOMEDTREE option in the PROC TREESPLIT statement to request diagrams that begin with other nodes and have specified depths.
The table in Figure 7 displays fit statistics for the tree model.
Figure 7: Fit Statistics
| Fit Statistics for Selected Tree | ||
|---|---|---|
| Number of Leaves | Misclassification Rate | |
| Training | 3 | 0.1124 |
The misclassification rate is the total proportion of the 178 wine samples that were misclassified. The following numbers of samples were misclassified in the terminal nodes:
node 3: 0
node 4: 1
node 5: 1
node 6: 4
So the total misclassification rate is .
[5] Disclaimer: SAS may reference other websites or content or resources for use at Customer's sole discretion. SAS has no control over any websites or resources that are provided by companies or persons other than SAS. Customer acknowledges and agrees that SAS is not responsible for the availability or use of any such external sites or resources, and does not endorse any advertising, products, or other materials on or available from such websites or resources. Customer acknowledges and agrees that SAS is not liable for any loss or damage that may be incurred by Customer or its end users as a result of the availability or use of those external sites or resources, or as a result of any reliance placed by Customer or its end users on the completeness, accuracy, or existence of any advertising, products, or other materials on, or available from, such websites or resources.