The KCLUS Procedure

Example 9.1 Cluster Analysis

This example uses the Iris data set in the Sashelp library to demonstrate how to use PROC KCLUS to perform cluster analysis. The iris data published by Fisher (1936) have been widely used for examples in discriminant and cluster analyses. The sepal length, sepal width, petal length, and petal width are measured in millimeters on 50 iris specimens from each of three species: Iris setosa, I. versicolor, and I. virginica. Mezzich and Solomon (1980) discuss a variety of cluster analyses that use the Iris data.

You can load the Sashelp.Iris data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step:

data mycas.iris;
   set sashelp.iris;
run;

These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.

The following statements perform clustering:

proc kclus
   data=mycas.iris
   seed=12345
   maxclusters=3
   outstat(outiter)=mycas.kclusOutstat1;
   input SepalLength SepalWidth PetalLength PetalWidth;
   score out=mycas.kclusOut1
   copyvars=(SepalLength SepalWidth PetalLength PetalWidth Species);
run;

In this example, PROC KCLUS generates the data table mycas.kclusOut1, which contains the cluster membership information for each observation in the input data table. For each observation, the mycas.kclusOut1 data table includes the variables that are specified in the COPYVARS= option in the SCORE statement and two new variables: _CLUSTER_ID_, which is the ID of the closest cluster, and _DISTANCE_, which is the distance between the observation and the centroid of the closest cluster. This example uses the variables in both the INPUT statement and the COPYVARS= option in order to transfer these variables to the output data table to do further analysis.

PROC KCLUS generates several ODS tables, some of which are shown in Figure 7 through Figure 12.

Figure 7: Number of Observations

The KCLUS Procedure

Number of Observations Read150
Number of Observations Used150


Figure 8: Model Information

Model Information
Clustering AlgorithmK-means
Maximum Iterations10
Stop CriterionCluster Change
Stop Criterion Value0
Clusters3
InitializationForgy
Seed12345
Distance for Interval VariablesEuclidean
StandardizationNone
Interval ImputationNone


Figure 9: Cluster Summary

Cluster Summary for Interval Variables
ClusterFrequencyDistance from Cluster Centroid
to Observation
SSEStandard
Deviation
Nearest
Cluster
Distance
to
Nearest
Cluster
Centroid
MinimumMaximumAverage
1392.394515.51567.31852541.48.0724217.8842
2612.357116.46807.31113829.17.9229117.8842
3500.661812.48034.81711515.15.5047233.4949


Figure 10: Iteration History

Iteration History
Iteration
Number
SSESSE ChangeStop
Criterion
071498  
113148-5835018.000000
28123.352556-5024.4905064.666667
37987.357983-135.9945732.000000
47934.436415-52.9215692.000000
57892.130972-42.3054420.666667
67885.566583-6.5643900


Figure 11: Descriptive Statistics

Descriptive Statistics
VariableMeanStandard
Deviation
SepalLength58.4333338.280661
SepalWidth30.5733334.358663
PetalLength37.58000017.652982
PetalWidth11.9933337.622377


Figure 12: Within-Cluster Statistics

Within Cluster Statistics
VariableClusterMeanStandard
Deviation
SepalLength168.53854.8820
 258.83614.4803
 350.06003.5249
SepalWidth130.76922.8696
 227.40982.9290
 334.28003.7906
PetalLength157.15385.1018
 243.88525.1157
 314.62001.7366
PetalWidth120.53852.9633
 214.34432.9994
 32.46001.0539


The following statements extract the first 10 observations from the output data table; they are shown in Figure 13.


proc print noobs data=mycas.kclusOut1(obs=10);
run;

Figure 13: First 10 Observations in the Output Data Table

SepalLengthSepalWidthPetalLengthPetalWidthSpecies_CLUSTER_ID__DISTANCE_
5033142Setosa31.4959946524
5133175Setosa33.8259639308
5234142Setosa32.1066561181
5035166Setosa33.8675573687
4830143Setosa34.8205808779
5030162Setosa34.5208406298
5840122Setosa310.140907257
5135142Setosa31.4135062787
5744154Setosa312.048153385
5241151Setosa37.1552777724


PROC KCLUS creates the output statistics data table, which contains the cluster centroids. This data table includes the iteration number as _ITERATION_, the cluster ID as _CLUSTER_ID_, and the cluster centroids, which consist of the variables that are specified in the INPUT statement. Because the OUTITER= suboption is included in the OUTSTAT= option in the PROC KCLUS statement, cluster centroids for each iteration are added to the kclusOutstat1 data table.

The following statements extract the centroids before the first iteration and after the last iteration:

proc print noobs data=mycas.kclusOutstat1(firstobs=1 obs=3);
run;

proc print noobs data=mycas.kclusOutstat1(firstobs=16 obs=18);
run;

Figure 14 and Figure 15 show the results.

Figure 14: Cluster Centroids before the First Iteration

_ITERATION__CLUSTER_ID_SepalLengthSepalWidthPetalLengthPetalWidth
0163254915
0261284712
0364294313


Figure 15: Cluster Centroids after the Last Iteration

_ITERATION__CLUSTER_ID_SepalLengthSepalWidthPetalLengthPetalWidth
5168.27530.7057.000020.6250
5258.85027.4043.766714.1833
5350.06034.2814.62002.4600


Last updated: April 08, 2021