The KCLUS Procedure
Example 9.1 Cluster Analysis
This example uses the Iris data set in the Sashelp library to demonstrate how to use PROC KCLUS to perform cluster analysis. The iris data published by Fisher (1936) have been widely used for examples in discriminant and cluster analyses. The sepal length, sepal width, petal length, and petal width are measured in millimeters on 50 iris specimens from each of three species: Iris setosa, I. versicolor, and I. virginica. Mezzich and Solomon (1980) discuss a variety of cluster analyses that use the Iris data.
You can load the Sashelp.Iris data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step:
data mycas.iris;
set sashelp.iris;
run;
These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.
The following statements perform clustering:
proc kclus
data=mycas.iris
seed=12345
maxclusters=3
outstat(outiter)=mycas.kclusOutstat1;
input SepalLength SepalWidth PetalLength PetalWidth;
score out=mycas.kclusOut1
copyvars=(SepalLength SepalWidth PetalLength PetalWidth Species);
run;
In this example, PROC KCLUS generates the data table mycas.kclusOut1, which contains the cluster membership information for each observation in the input data table. For each observation, the mycas.kclusOut1 data table includes the variables that are specified in the COPYVARS= option in the SCORE statement and two new variables: _CLUSTER_ID_, which is the ID of the closest cluster, and _DISTANCE_, which is the distance between the observation and the centroid of the closest cluster. This example uses the variables in both the INPUT statement and the COPYVARS= option in order to transfer these variables to the output data table to do further analysis.
PROC KCLUS generates several ODS tables, some of which are shown in Figure 7 through Figure 12.
Figure 7: Number of Observations
| Number of Observations Read | 150 |
|---|---|
| Number of Observations Used | 150 |
Figure 8: Model Information
| Model Information | |
|---|---|
| Clustering Algorithm | K-means |
| Maximum Iterations | 10 |
| Stop Criterion | Cluster Change |
| Stop Criterion Value | 0 |
| Clusters | 3 |
| Initialization | Forgy |
| Seed | 12345 |
| Distance for Interval Variables | Euclidean |
| Standardization | None |
| Interval Imputation | None |
Figure 9: Cluster Summary
| Cluster Summary for Interval Variables | ||||||||
|---|---|---|---|---|---|---|---|---|
| Cluster | Frequency | Distance from Cluster Centroid to Observation | SSE | Standard Deviation | Nearest Cluster | Distance to Nearest Cluster Centroid | ||
| Minimum | Maximum | Average | ||||||
| 1 | 39 | 2.3945 | 15.5156 | 7.3185 | 2541.4 | 8.0724 | 2 | 17.8842 |
| 2 | 61 | 2.3571 | 16.4680 | 7.3111 | 3829.1 | 7.9229 | 1 | 17.8842 |
| 3 | 50 | 0.6618 | 12.4803 | 4.8171 | 1515.1 | 5.5047 | 2 | 33.4949 |
Figure 10: Iteration History
| Iteration History | |||
|---|---|---|---|
| Iteration Number | SSE | SSE Change | Stop Criterion |
| 0 | 71498 | ||
| 1 | 13148 | -58350 | 18.000000 |
| 2 | 8123.352556 | -5024.490506 | 4.666667 |
| 3 | 7987.357983 | -135.994573 | 2.000000 |
| 4 | 7934.436415 | -52.921569 | 2.000000 |
| 5 | 7892.130972 | -42.305442 | 0.666667 |
| 6 | 7885.566583 | -6.564390 | 0 |
Figure 11: Descriptive Statistics
| Descriptive Statistics | ||
|---|---|---|
| Variable | Mean | Standard Deviation |
| SepalLength | 58.433333 | 8.280661 |
| SepalWidth | 30.573333 | 4.358663 |
| PetalLength | 37.580000 | 17.652982 |
| PetalWidth | 11.993333 | 7.622377 |
Figure 12: Within-Cluster Statistics
| Within Cluster Statistics | |||
|---|---|---|---|
| Variable | Cluster | Mean | Standard Deviation |
| SepalLength | 1 | 68.5385 | 4.8820 |
| 2 | 58.8361 | 4.4803 | |
| 3 | 50.0600 | 3.5249 | |
| SepalWidth | 1 | 30.7692 | 2.8696 |
| 2 | 27.4098 | 2.9290 | |
| 3 | 34.2800 | 3.7906 | |
| PetalLength | 1 | 57.1538 | 5.1018 |
| 2 | 43.8852 | 5.1157 | |
| 3 | 14.6200 | 1.7366 | |
| PetalWidth | 1 | 20.5385 | 2.9633 |
| 2 | 14.3443 | 2.9994 | |
| 3 | 2.4600 | 1.0539 | |
The following statements extract the first 10 observations from the output data table; they are shown in Figure 13.
proc print noobs data=mycas.kclusOut1(obs=10);
run;
Figure 13: First 10 Observations in the Output Data Table
| SepalLength | SepalWidth | PetalLength | PetalWidth | Species | _CLUSTER_ID_ | _DISTANCE_ |
|---|---|---|---|---|---|---|
| 50 | 33 | 14 | 2 | Setosa | 3 | 1.4959946524 |
| 51 | 33 | 17 | 5 | Setosa | 3 | 3.8259639308 |
| 52 | 34 | 14 | 2 | Setosa | 3 | 2.1066561181 |
| 50 | 35 | 16 | 6 | Setosa | 3 | 3.8675573687 |
| 48 | 30 | 14 | 3 | Setosa | 3 | 4.8205808779 |
| 50 | 30 | 16 | 2 | Setosa | 3 | 4.5208406298 |
| 58 | 40 | 12 | 2 | Setosa | 3 | 10.140907257 |
| 51 | 35 | 14 | 2 | Setosa | 3 | 1.4135062787 |
| 57 | 44 | 15 | 4 | Setosa | 3 | 12.048153385 |
| 52 | 41 | 15 | 1 | Setosa | 3 | 7.1552777724 |
PROC KCLUS creates the output statistics data table, which contains the cluster centroids. This data table includes the iteration number as _ITERATION_, the cluster ID as _CLUSTER_ID_, and the cluster centroids, which consist of the variables that are specified in the INPUT statement. Because the OUTITER= suboption is included in the OUTSTAT= option in the PROC KCLUS statement, cluster centroids for each iteration are added to the kclusOutstat1 data table.
The following statements extract the centroids before the first iteration and after the last iteration:
proc print noobs data=mycas.kclusOutstat1(firstobs=1 obs=3);
run;
proc print noobs data=mycas.kclusOutstat1(firstobs=16 obs=18);
run;
Figure 14 and Figure 15 show the results.
Figure 14: Cluster Centroids before the First Iteration
| _ITERATION_ | _CLUSTER_ID_ | SepalLength | SepalWidth | PetalLength | PetalWidth |
|---|---|---|---|---|---|
| 0 | 1 | 63 | 25 | 49 | 15 |
| 0 | 2 | 61 | 28 | 47 | 12 |
| 0 | 3 | 64 | 29 | 43 | 13 |
Figure 15: Cluster Centroids after the Last Iteration
| _ITERATION_ | _CLUSTER_ID_ | SepalLength | SepalWidth | PetalLength | PetalWidth |
|---|---|---|---|---|---|
| 5 | 1 | 68.275 | 30.70 | 57.0000 | 20.6250 |
| 5 | 2 | 58.850 | 27.40 | 43.7667 | 14.1833 |
| 5 | 3 | 50.060 | 34.28 | 14.6200 | 2.4600 |